Files
msd-core/src/phase-estimation.cts
Tom Boucher 0624c5da6f chore(#3212): src/text-lines.cts is the sole owner of line-terminator handling — Phase 2 (#3420)
* test(#3413): failing-first suite for the line-terminator seam

Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Tests only — src/text-lines.cts
does not exist yet, so tests/text-lines.test.cjs fails with MODULE_NOT_FOUND
at its require line, which is the intended RED.

The frontmatter.test.cjs additions drive #3360 (confirmed-bug) fail-first:
parseMustHavesBlock currently returns [] for every must_haves block on a
CRLF-authored plan file, because \r is its own LineTerminator in ECMAScript
and two /m-anchored \s* patterns can absorb it, inflating a captured indent
by one character and tripping the "not nested under must_haves" guard.
Verified locally against the current (unfixed) compiled module: both the
direct repro and the silent-exit "blank line before must_haves:" variant
return [] today. A parity property test (crlf vs lf must deep-equal for
every block name) matches a pattern this maintainer has required repeatedly
for prior CRLF fixes in this codebase (Cortex-recorded, verify_intent=held).

The no-crlf-fragile-split.rule.test.cjs additions lock the eslint rule's
future fix-hint text (pointing at splitLines()) and its self-reference
non-violation (the seam's own correct \r?\n split must never flag itself).

Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md
Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md

* chore(#3413): src/text-lines.cts owns line-terminator handling

Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Adds splitLines/normalizeEol/
detectEol/joinLines and migrates frontmatter.cts onto it.

parseMustHavesBlock (#3360, confirmed-bug) returned [] for every
must_haves block on a CRLF plan file. Root cause: \r is its own
LineTerminator in ECMAScript, so under /m two \s*-anchored indentation
lookups could match at the position INSIDE a \r\n pair and absorb the
terminator, inflating the captured indent by one character and tripping
the "not nested under must_haves" guard. Two silent exits, one with a
diagnostic and one without (a blank line before must_haves: hits the
silent path). Fixed by converting both lookups from a whole-string /m
match to split-then-scan — splitLines first, then a per-line, non-/m
match — the same structural pattern parseYamlRegion (30 lines away in
the same file) already used safely. Nothing downstream of the two
lookups changed; blockLines is now sliced from the already-split array
instead of re-splitting a substring, but its contents are unchanged for
LF input, and the per-line dash/kv parsing loop is untouched.

A parity property test (CRLF and LF plans parse to identical must_haves
for every block name) matches a pattern this maintainer has required
repeatedly for prior CRLF fixes in this file's neighborhood (Cortex:
7 recorded decisions, verify_intent -> held).

frontmatter.cts's other .split(/\r?\n/) call sites (parseYamlRegion,
isFrontmatterShaped, sliceTopLevelFrontmatterSegments, spliceFrontmatter)
are rerouted onto splitLines — a literal 1:1 substitution, zero behavior
change, since splitLines IS that same regex plus a type guard.

The 4 scripts/normalizeLineEndings copies (gen-registry, gen-loop-host-
contract, gen-capability-registry, gen-context-index) are deleted and
rerouted onto normalizeEol, which strips a bare unpaired \r exactly like
the deleted copies did (not just \r\n pairs) -- verified against each
script's own --check mode against its real generated output.

local/no-crlf-fragile-split widens from tests/ to src/**/*.cts, with its
fix-hint message now naming splitLines() instead of the raw regex --
the prohibition finally has a primitive to point at. Detection logic
unchanged in this phase (deliberate scope limit, see design doc Known
limits: the rule doesn't yet recognize safeReadFile/platformReadSync as
a content source, and has no detector for the \s-adjacent-to-anchor
shape that is #3360's actual mechanism -- the CLASS is converged by the
direct fix + regression test regardless).

joinLines/detectEol are NOT wired into frontmatter.cts's own write path
(cmdFrontmatterSet/Merge -> platformWriteSync) -- verified that
platformWriteSync already, unconditionally converts CRLF->LF on every
.md write today as a pre-existing policy owned by a different module,
and ADR-3212's backward-compatibility clause rules out a file-format
change in any phase. Stated explicitly in Known limits rather than left
for a reader to discover.

Six-gate ripple: .gitignore, eslint.config.mjs (src/**/*.cts block),
docs/INVENTORY.md + INVENTORY-MANIFEST.json (regenerated), CONTEXT.md
glossary (Text Lines Module, mirroring Phase 1's Pattern Module entry).

Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md
Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md

* fix(#3413): fix 13 pre-existing CRLF-fragile splits the widened rule found

Widening local/no-crlf-fragile-split from tests/ to src/**/*.cts (the
previous commit) immediately surfaced 13 real, pre-existing violations
across 10 files -- undetected until now because the rule never scanned
src/. This is the exact defect class ADR-3212 exists to close, playing
out again one phase after Phase 1 hit the same shape ("the new lint
rule -- once live -- found 27 more"). Per CLAUDE.md's no-defer rule,
fixed inline rather than deferred or suppressed; there is no
established suppression convention for this rule in src/ and inventing
one now would undermine the point of widening it.

audit.cts, broken-windows.cts, core-utils.cts, init.cts, milestone.cts,
phase.cts (x3), profile-output.cts, roadmap.cts (x2): bare-\n splits or
regex character classes widened to \r?\n / [^\r\n], each following the
same pattern already established migrating frontmatter.cts.

phase-estimation.cts: `\r?(?:\n|$)` restructured to `(?:\r?\n|\r?$)` --
already semantically CRLF-safe, but the rule's lexical scanner doesn't
recognize \r? guarding a group (only \r? immediately before a literal
\n). Verified the two forms are equivalent across all four EOL/EOF
cases before restructuring, not assumed.

roadmap-upgrade.cts needed two coupled sites, not the one flagged line:
computeMigrationPlan and applyMigration must agree on line
representation for the lines[edit.lineIndex] === edit.from equality
check to hold, and the write-back needed joinLines + detectEol -- a
plain lines.join('\n') was silently flattening a CRLF ROADMAP.md to LF
wholesale on every migration. This is the first real production
consumer of joinLines/detectEol in this epic (frontmatter.cts's own
write path doesn't use them -- see the previous commit's Known limits).

Fixing the 13 flagged sites surfaced 4 more adjacent same-shape sites
the rule doesn't track (.search() and new RegExp(dynamicString) aren't
in its tracked call/construction set). Investigated each empirically --
hand-tracing this exact bug class already produced one wrong conclusion
earlier in this phase (a detectEol design-doc arithmetic error), so
these were verified with real CRLF fixtures rather than reasoned about
on paper:

  - audit.cts (scanTodos): REAL bug, fixed. `bodyMatch.trim().split
    ('\n')[0]` leaked a trailing \r into a user-visible todo summary on
    CRLF input -- .trim() only strips the string's outer edges, not a
    \r sitting mid-string before the first bare \n. Now splitLines(...)
    [0].
  - phase.cts (cmdPhaseInsert, bullet-style branch): REAL bug, fixed.
    [^\n]* in targetBulletPattern swallowed a line's trailing \r on
    CRLF input, shifting the computed insert position to land INSIDE
    the \r\n pair; combined with a hardcoded '\n' bullet separator, a
    CRLF ROADMAP.md ended up with a mixed CRLF/LF result after an
    insert. Fixed with two coupled changes (either alone still
    corrupts, verified both ways): [^\r\n]* in the pattern, and the new
    bullet's leading terminator now comes from detectEol(rawContent).
  - roadmap.cts (cmdRoadmapAnnotateDependencies phase-boundary scan):
    investigated, genuinely safe, left untouched. The .search(/\n#{2,4}
    .../) boundary-finder and the [^\n]*-based heading match were
    empirically verified on a 3-phase CRLF fixture -- the only stray \r
    ends up at the tail of an intermediate phaseSection string that is
    only ever used for .test()-based idempotency checks, never for an
    exact-match comparison or written back to disk. No corruption on
    round-trip.

Every fix re-verified: npm run build:lib clean, npx eslint
'src/**/*.cts' --no-cache reports 0 problems (was 13), and each
fixed function's existing LF-input tests were spot-checked unchanged.

* fix(#3413): apply orthogonal review findings

Two isolated review engines (correctness + security) ran against the
full diff and found three majors, one real security issue, and several
disclosure-worthy minors. All fixed or explicitly disclosed with
evidence; nothing deferred.

MAJOR — detectEol's tie-break contradicted its own documented contract.
Code returned '\n' on a 1:1 crlf/bare-LF tie; every doc (design doc,
CONTEXT.md, the function's own comment) says ties resolve to '\r\n'.
The existing test masked this by reusing the same tie fixture the
buggy code happened to satisfy, rather than a genuine LF-majority
case. Root cause: an Edit attempted earlier in this phase to fix this
exact arithmetic error was blocked by the tier guard, and a later
dispatch was incorrectly told it had already landed. Fixed: condition
is now crlfCount >= bareLfCount; the test fixture corrected to a
genuine 2:1 majority, with a new explicit tie-case test.

MAJOR — phase.cts's cmdPhaseInsert built an EOL-aware bulletEntry via
detectEol(rawContent), justified by a comment claiming a hardcoded
'\n' corrupts a CRLF ROADMAP.md. False: this write goes through
platformWriteSync, whose normalizeContent/_normalizeMd unconditionally
converts CRLF->LF for any .md target — the templating was inert dead
code, erased before the file is ever written. Reverted to hardcoded
'\n', comment corrected to state the true reasoning. The separate
[^\n]* -> [^\r\n]* widening one function up (a real splice-position
fix, independent of final EOL) was kept.

MAJOR — roadmap-upgrade.cts's stated rationale for switching onto
splitLines/joinLines was wrong (both functions always agreed on line
representation, before and after — the claimed equality-check risk
never existed), and the change it justified introduced a real
regression: forcing every line onto one dominant terminator silently
rewrites untouched lines' EOL on a mixed-CRLF/LF ROADMAP.md. This
write path uses raw fs.writeFileSync, not platformWriteSync, so unlike
the phase.cts case above the regression is genuinely live.

Fixing this took two attempts. The first attempt (revert to
split('\n')/join('\n') plus a suppression comment) was correctly
blocked by an agent that discovered local/no-crlf-fragile-split is a
PROTECTED_RULES entry in tests/portability-rule-disable-ban.test.cjs —
a hard, out-of-band, ADR-1703-governed guardrail banning any
eslint-disable of this rule anywhere in src/**/*.cts. That agent also
detected and correctly disregarded an injected instruction that
appeared in tool output during a git operation, per this session's
untrusted-content policy. The actual fix: computeMigrationPlan
reverted to roadmapContent.split('\n') (confirmed lint-clean — the
rule's data-flow tracking only follows a variable's initializer, and
this one is declared empty then reassigned in a try block).
applyMigration's write-back now splices edits against the ORIGINAL
content string via indexOf('\n', pos) boundary-walking instead of a
full split/rejoin, so every untouched character — including every
line's own terminator — is copied byte-for-byte. A capture-group split
(/(\r\n|\n)/, preserving terminators inline) was tried first and
empirically confirmed to still trip the rule before this approach was
chosen instead.

MINOR (security) — roadmap.cts's cmdRoadmapAnnotateDependencies used
the STRING form of String#replace, so $&, $`, $', $1-$9 inside
must_haves.truths content (author-controlled) were interpreted as
replacement directives, splicing unrelated ROADMAP.md text into the
result. Fixed with the function-replacement form, which is never
pattern-interpreted. Verified before/after with the reviewer's exact
repro.

Also disclosed rather than silently left: test matrix row 31 (four
planned CRLF-materialized regression tests) was never implemented as
separate files — corrected to record the actual verification (a
manual --check run plus incidental existing coverage via each script's
normalizeLineEndings: normalizeEol alias). parseMustHavesBlock's LF
behavior was claimed byte-for-byte unchanged but the old
yaml.indexOf(blockMatch[0]) substring search could match an unrelated
earlier occurrence of the header text (e.g. inside a quoted value) —
the split-then-scan fix incidentally also closes this, a strict
improvement now recorded in the design doc rather than left implicit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3413): checkpoint 2 red — missing eslint ignore entry, RuleTester config error

Checkpoint 2 came back red with 5 failures on the reviewed sha, both
gaps genuinely undetectable by any local gate.

eslint.config.mjs was missing the 'gsd-core/bin/lib/text-lines.cjs'
ignores-list entry (ADR-457: generated .cjs artifacts are excluded from
direct type-aware linting). Phase 1's sibling entry (pattern.cjs) sits
two lines above it and was the exact precedent read while researching
the six-gate ripple for this module -- missed anyway. Caught by
tests/repo-invariants.test.cjs's bin/lib coverage-tracking test, which
only runs on the remote suite.

tests/no-crlf-fragile-split.rule.test.cjs's row-32 case specified both
`messageId` and `message` on the same RuleTester error assertion --
ESLint's RuleTester rejects that combination outright. This existed
since the test was first authored and was never caught locally: `npx
eslint` only lints the file's syntax, it does not execute RuleTester,
and local `node --test` is hard-blocked in this repo -- the assertion
had never actually RUN before this checkpoint. It was even present in
checkpoint 1's failure list, listed there as one of the "expected RED"
tests; I matched it against my expected-failures list by test NAME
only and never inspected the actual failure detail closely enough to
notice it was failing for the wrong reason (a RuleTester config error,
not the intended message-text mismatch). Fixed by keeping `message`
(the exact-text assertion the test exists to make) and dropping
`messageId`. Verified the crlfFragileSplit message string in
eslint-rules/no-crlf-fragile-split.cjs matches this assertion
character-for-character, and swept every other invalid case in the
file for the same double-specification bug (none found -- all
pre-existing cases use messageId alone).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3413): add Fixed changeset for the #3360 CRLF parsing fix

The sole user-visible effect of this phase. No breaking-change label
or Changed fragment needed — ADR-3212's Backward Compatibility section
names the Node floor (Phase 1, already shipped) as the epic's only
breaking change; Phase 2 has none.

* chore(#3413): backfill changeset pr number to 3420

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 20:27:48 -04:00

516 lines
22 KiB
TypeScript

/**
* Phase Estimation — estimate/actuals schema, smart-zone threshold policy, and
* estimate-vs-actual calibration.
*
* Epic #1952, Phase 1 (#2630). Design lock: docs/adr/2629-phase-effort-estimation-calibration.md.
*
* Pure functions only — no I/O, no config reads. Callers supply the budget and
* the raw calibration document; this module decides policy over them. The CLI
* seam (gsd-tools) owns reading `.planning/config.json` and
* `.planning/estimation-calibration.json`.
*
* Two properties this module exists to preserve, both from ADR-2629:
*
* 1. Every signal is EXOGENOUS. The correction routes on a measured
* actual/estimate ratio; `confidence` routes on a calibration sample
* count. Nothing routes on a model's self-assessment. This project
* measured self-rated confidence and found it weak
* (gsd-core/references/honest-verifier.md:25-29 — "on a true blind spot it
* stays confidently wrong"), which is why deriveConfidence() takes a
* sample count and there is no "how sure are you?" input anywhere here.
*
* 2. Estimate and actual share ONE measurement scale — estimateTokens() from
* prompt-budget. A ratio between two different measurement methods would
* measure the methods, not the miss. measureTokens() below is the single
* re-export so no consumer reaches for a second estimator.
*
* ADR-457 build-at-publish: source here, compiled to
* gsd-core/bin/lib/phase-estimation.cjs (gitignored).
*/
// eslint-disable-next-line @typescript-eslint/no-require-imports -- prompt-budget.cjs is an export= CommonJS module
import promptBudget = require('./prompt-budget.cjs');
const { estimateTokens } = promptBudget;
/** Confidence in an estimate. DERIVED from calibration sample count — never self-rated. */
export type Confidence = 'low' | 'med' | 'high';
declare const RAW_TOKENS_BRAND: unique symbol;
declare const CALIBRATED_TOKENS_BRAND: unique symbol;
/**
* A token count the correction factor has NOT been applied to — the planner's
* uncorrected projection, and the only legal denominator for calibration.
*
* Both types below are compile-time brands (#2671). They erase entirely: the
* emitted `.cjs` sees plain numbers, the CLI's JSON output is unchanged, and
* every untyped `.cjs` caller keeps working exactly as before. What they buy is
* that the two states stop being interchangeable `number`s at the seams where
* epic #1952 twice mixed them up:
*
* - #2631 — an already-calibrated figure was fed to a parameter named
* `rawTokens`, so the correction became factor^2 (4x under to 9x over under
* the [0.5, 3.0] clamp), invisible below 3 samples because factor === 1
* there and 1^2 === 1.
* - #2632 — calibration measured actual/calibrated instead of actual/raw, so
* the loop un-corrected itself and settled near 1.41 instead of 2.0.
*
* Both were composition errors between individually-correct functions, and both
* shipped past a green ~26,800-test suite. `--calibrated` and `raw_tokens` fix
* the two known call sites but remain conventions a caller must remember; the
* brands make the wrong composition unrepresentable instead. Compile fixtures:
* `tests/fixtures/brand-typing/`.
*/
export type RawTokens = number & { readonly [RAW_TOKENS_BRAND]: true };
/**
* A token count the correction factor HAS been applied to — what a plan records
* in `estimate.tokens` and the only figure meaningful against the smart-zone
* budget.
*
* "Applied" is about provenance, not arithmetic: a project below
* MIN_CALIBRATION_SAMPLES has factor 1, so the calibrated figure equals the raw
* one numerically while still being a different thing to a reader and to the
* calibration loop.
*/
export type CalibratedTokens = number & { readonly [CALIBRATED_TOKENS_BRAND]: true };
/**
* Assert that a bare number is an UNCORRECTED projection.
*
* Call this only where a number crosses a trust boundary carrying a basis the
* type system cannot see — argv, disk frontmatter, a persisted document. The
* parameter type refuses a `CalibratedTokens`, so a corrected figure cannot be
* laundered back into the basis; without that the brand would be decorative and
* #2632 would be one keystroke away again.
*/
export function asRawTokens(tokens: number & { readonly [CALIBRATED_TOKENS_BRAND]?: never }): RawTokens {
return tokens as RawTokens;
}
/** Assert that a bare number already has the correction applied. Refuses a `RawTokens`. */
export function asCalibratedTokens(tokens: number & { readonly [RAW_TOKENS_BRAND]?: never }): CalibratedTokens {
return tokens as CalibratedTokens;
}
export const CONFIDENCE_VALUES: readonly Confidence[] = Object.freeze(['low', 'med', 'high'] as const);
/** Below this many calibration samples, no correction is applied (ADR-2629 Decision 4). */
export const MIN_CALIBRATION_SAMPLES = 3;
/** Sample-count thresholds for derived confidence (ADR-2629 Decision 1). */
export const CONFIDENCE_MED_MIN_SAMPLES = 3;
export const CONFIDENCE_HIGH_MIN_SAMPLES = 6;
/** Correction-factor clamp. Outside this range the estimator is wrong in kind, not degree. */
export const CALIBRATION_FACTOR_MIN = 0.5;
export const CALIBRATION_FACTOR_MAX = 3.0;
/** Schema version for the persisted calibration document. */
export const CALIBRATION_SCHEMA_VERSION = 1;
export interface PhaseEstimate {
/** Calibrated at emission time per ADR-2629 Decision 1 — never the raw projection. */
tokens: CalibratedTokens;
tasks: number;
confidence: Confidence;
/**
* The planner's UNCALIBRATED projection, before the correction factor was
* applied. Optional for backward compatibility with plans written before
* #2632.
*
* Calibration MUST measure actual/raw, not actual/calibrated. Measuring
* against the already-corrected figure makes the loop self-defeating: once
* the correction works, the observed ratio approaches 1, which drags the
* median back toward 1, which un-corrects the next estimate. Simulated over
* 10 phases with a true 2x underestimate, that oscillates and settles at
* ~1.41 instead of converging on 2.0.
*/
rawTokens?: RawTokens;
}
export interface PhaseActuals {
tokens: number;
tasks: number;
commits: number;
}
export interface BudgetClassification {
/** True only when the estimate strictly exceeds the budget. At the budget exactly, false. */
overBudget: boolean;
/** estimate / budget. 0 when the budget is unusable. */
ratio: number;
/** Human-facing split advice. Null unless overBudget. Advisory — never a block. */
recommendation: string | null;
/** False when the supplied budget was not a positive finite number. */
budgetValid: boolean;
}
export interface CalibrationSample {
/**
* The RAW basis — ADR-2629 Decision 4. Branded so a calibrated figure cannot
* take this slot: that substitution is #2632, and it is silent at runtime
* because both sides are positive integers of the same magnitude.
*/
estimateTokens: RawTokens;
/**
* Measured cost on the `estimateTokens` scale. Deliberately unbranded — an
* actual is neither a projection nor a correction of one, so giving it either
* brand would make the type say something untrue.
*/
actualTokens: number;
}
export interface CalibrationResult {
/** Multiply a raw estimate by this. Exactly 1 when not applied. */
factor: number;
/** Count of USABLE samples (both sides positive and finite). */
sampleCount: number;
/** True once sampleCount >= MIN_CALIBRATION_SAMPLES. */
applied: boolean;
/** Derived from sampleCount — the same signal, surfaced for the estimate block. */
confidence: Confidence;
/** True when the median ratio fell outside the clamp and was pinned to a bound. */
clamped: boolean;
}
/**
* A positive, finite, safe integer. Rejects NaN, Infinity, negatives, zero,
* non-integers, and anything past MAX_SAFE_INTEGER (where integer arithmetic
* silently stops being exact).
*/
function isPositiveInt(value: unknown): value is number {
return typeof value === 'number'
&& Number.isSafeInteger(value)
&& value > 0;
}
function isPositiveFinite(value: unknown): value is number {
return typeof value === 'number' && Number.isFinite(value) && value > 0;
}
function isConfidence(value: unknown): value is Confidence {
return typeof value === 'string' && (CONFIDENCE_VALUES as readonly string[]).includes(value);
}
/**
* A usable calibration sample: both sides present, positive, and finite.
* A zero or negative estimate would divide to Infinity or flip the ratio's
* sign, so those are dropped rather than coerced.
*/
function isCalibrationSample(value: unknown): value is CalibrationSample {
if (value === null || typeof value !== 'object' || Array.isArray(value)) return false;
const record = value as Record<string, unknown>;
// The RawTokens brand on estimateTokens is asserted here, at the disk trust
// boundary — a persisted sample's basis is a fact about the writer, and the
// only writers are collectCalibrationSamples() (which reads it through
// calibrationBasis()) and this module's own renderCalibrationDocument().
return isPositiveFinite(record['estimateTokens']) && isPositiveFinite(record['actualTokens']);
}
/**
* Measure text on the canonical scale. The ONE estimator both the estimate and
* the actuals must use — see property 2 in the module header.
*/
export function measureTokens(text: string | null | undefined): number {
return estimateTokens(text);
}
/**
* Derive confidence from how much measured history backs the estimate.
*
* Exogenous by construction: the input is a count, not a judgment. A non-integer
* or negative count degrades to 'low' rather than throwing — an unusable history
* is exactly the low-confidence case.
*/
export function deriveConfidence(sampleCount: unknown): Confidence {
if (typeof sampleCount !== 'number' || !Number.isFinite(sampleCount) || sampleCount < 0) return 'low';
if (sampleCount >= CONFIDENCE_HIGH_MIN_SAMPLES) return 'high';
if (sampleCount >= CONFIDENCE_MED_MIN_SAMPLES) return 'med';
return 'low';
}
/**
* Classify an estimate against the smart-zone budget.
*
* Boundary contract (ADR-2629 Decision 3 + RULESET.TESTS.boundary-coverage.fixtures):
* budget-1 → under, budget → under, budget+1 → over. The comparison is strictly
* greater-than, so landing exactly on the budget is not a violation.
*
* An unusable budget (hand-edited config, missing key) never fabricates a
* violation: it reports budgetValid=false and overBudget=false, so a broken
* config cannot spam split recommendations.
*/
export function classifyAgainstBudget(estimate: CalibratedTokens, budget: number): BudgetClassification {
// Kept for untyped `.cjs` callers — see the note in applyCalibration. A
// hand-edited config reaches `budget` as anything at runtime regardless of
// what the TypeScript signature promises.
if (!isPositiveFinite(budget) || !isPositiveFinite(estimate)) {
return { overBudget: false, ratio: 0, recommendation: null, budgetValid: isPositiveFinite(budget) };
}
const ratio = estimate / budget;
if (estimate <= budget) {
return { overBudget: false, ratio, recommendation: null, budgetValid: true };
}
const slices = Math.ceil(ratio);
return {
overBudget: true,
ratio,
recommendation:
`Estimated ${estimate} tokens exceeds the ${budget}-token smart-zone budget `
+ `(${ratio.toFixed(2)}x). Consider splitting this phase into about ${slices} `
+ `slices — a tracer plus ${slices - 1} expansion slice(s) — so each runs inside the budget.`,
budgetValid: true,
};
}
/** Median of a non-empty numeric array. Caller guarantees non-empty. */
function median(sorted: number[]): number {
const mid = Math.floor(sorted.length / 2);
if (sorted.length % 2 === 1) return sorted[mid];
return (sorted[mid - 1] + sorted[mid]) / 2;
}
/**
* Compute the correction factor from estimate/actual history.
*
* Median, not mean — one pathological phase (an aborted run, a mass rename)
* must not swing every later projection. Clamped, because a ratio outside
* [0.5, 3.0] means the estimator is wrong in kind and amplifying it would make
* the next estimate worse, not better.
*
* Samples missing either side, or carrying a non-positive/non-finite value, are
* dropped rather than coerced — a zero estimate would divide to Infinity.
*/
export function computeCalibration(samples: unknown): CalibrationResult {
const candidates: unknown[] = Array.isArray(samples) ? samples : [];
const usable = candidates.filter(isCalibrationSample);
const sampleCount = usable.length;
const confidence = deriveConfidence(sampleCount);
if (sampleCount < MIN_CALIBRATION_SAMPLES) {
return { factor: 1, sampleCount, applied: false, confidence, clamped: false };
}
const ratios = usable.map((s) => s.actualTokens / s.estimateTokens).sort((a, b) => a - b);
const raw = median(ratios);
const factor = Math.min(CALIBRATION_FACTOR_MAX, Math.max(CALIBRATION_FACTOR_MIN, raw));
return { factor, sampleCount, applied: true, confidence, clamped: factor !== raw };
}
/**
* Apply a correction factor to a raw estimate. Rounds to an integer because
* `estimate.tokens` is an integer field; floors at 1 so a heavy shrink factor
* can never produce a zero-token estimate.
*/
export function applyCalibration(rawTokens: RawTokens, factor: number): CalibratedTokens {
// These two guards look dead to the type-checker and are not: this module is
// compiled to `.cjs` and consumed by untyped callers (gsd-tools.cjs, the test
// suite), which reach it with NaN, null, 0 and worse. The brands are a
// compile-time contract for TypeScript callers; validation is what defends
// everyone else. Do not delete either one because the parameter is now typed.
if (!isPositiveFinite(rawTokens)) return asCalibratedTokens(0);
if (!isPositiveFinite(factor)) return asCalibratedTokens(Math.max(1, Math.round(rawTokens)));
// Bound the product: an inexact float past MAX_SAFE_INTEGER would masquerade
// as an integer token count. Unreachable through today's CLI (which is
// safe-integer bounded) but the function is exported and must not depend on
// its caller for that guarantee.
const scaled = Math.round(rawTokens * factor);
return asCalibratedTokens(Math.min(Number.MAX_SAFE_INTEGER, Math.max(1, scaled)));
}
/**
* Extract a two-space-indented scalar block (`estimate:` / `actuals:`) out of a
* document's leading YAML frontmatter.
*
* Hand-rolled because gsd-core ships no external dependencies (CONTRIBUTING.md
* "No external dependencies in core") — js-yaml is a devDependency and is not
* available at runtime. Scope is deliberately narrow: the leading `---` block
* only, so a `estimate:` line inside a fenced code block in the body cannot be
* mistaken for frontmatter (the DEFECT.FRONTMATTER-SCALAR-BROAD-GREP class).
*
* Numeric-looking values are returned as numbers so parseEstimate/parseActuals
* see the types they validate; everything else stays a string.
*/
export function extractFrontmatterBlock(text: unknown, key: string): Record<string, unknown> | null {
if (typeof text !== 'string') return null;
// Anchor at byte 0 — CRLF-tolerant.
const fm = /^---\r?\n([\s\S]*?)\r?\n---(?:\r?\n|\r?$)/.exec(text);
if (fm === null) return null;
const lines = fm[1].split(/\r?\n/);
const startIdx = lines.findIndex((l) => l === `${key}:` || l.startsWith(`${key}:`));
if (startIdx === -1) return null;
const out: Record<string, unknown> = Object.create(null) as Record<string, unknown>;
for (let i = startIdx + 1; i < lines.length; i += 1) {
const line = lines[i];
if (!/^\s/.test(line)) break; // dedent ends the block
const m = /^\s+([A-Za-z_][\w-]*):\s*(.*)$/.exec(line);
if (m === null) continue;
const rawValue = m[2].replace(/\s+#.*$/, '').trim();
if (rawValue === '') continue;
const asNumber = Number(rawValue);
out[m[1]] = /^-?\d+(?:\.\d+)?$/.test(rawValue) && Number.isFinite(asNumber)
? asNumber
: rawValue.replace(/^['"]|['"]$/g, '');
}
return Object.keys(out).length > 0 ? { ...out } : null;
}
/** Pull the `estimate:` mapping out of an already-parsed frontmatter object. */
function estimateBlockOf(input: unknown): unknown {
if (input === null || typeof input !== 'object') return null;
const record = input as Record<string, unknown>;
return Object.prototype.hasOwnProperty.call(record, 'estimate') ? record['estimate'] : record;
}
/**
* Parse an estimate block. Returns null for anything that is not a complete,
* well-typed estimate — a partial block is not a usable estimate, and silently
* defaulting a missing field would fabricate data the planner never produced.
*
* Accepts either the whole frontmatter object (`{estimate: {...}}`) or the
* estimate mapping itself, so callers need not unwrap.
*/
export function parseEstimate(input: unknown): PhaseEstimate | null {
const block = estimateBlockOf(input);
if (block === null || typeof block !== 'object' || Array.isArray(block)) return null;
const record = block as Record<string, unknown>;
const tokens = record['tokens'];
const tasks = record['tasks'];
const confidence = record['confidence'];
if (!isPositiveInt(tokens) || !isPositiveInt(tasks) || !isConfidence(confidence)) return null;
// The frontmatter trust boundary: `tokens` is calibrated-at-emission and
// `raw_tokens` is the uncorrected projection (ADR-2629 Decision 1/4), so this
// is where each figure's basis becomes a type rather than a field name.
const rawTokens = record['raw_tokens'];
return isPositiveInt(rawTokens)
? { tokens: asCalibratedTokens(tokens), tasks, confidence, rawTokens: asRawTokens(rawTokens) }
: { tokens: asCalibratedTokens(tokens), tasks, confidence };
}
/** Pull the `actuals:` mapping out of an already-parsed frontmatter object. */
function actualsBlockOf(input: unknown): unknown {
if (input === null || typeof input !== 'object') return null;
const record = input as Record<string, unknown>;
return Object.prototype.hasOwnProperty.call(record, 'actuals') ? record['actuals'] : record;
}
/**
* Parse an actuals block. `commits` may be 0 — a phase can legitimately record
* zero commits — so it is validated as a non-negative integer while tokens and
* tasks stay strictly positive.
*/
export function parseActuals(input: unknown): PhaseActuals | null {
const block = actualsBlockOf(input);
if (block === null || typeof block !== 'object' || Array.isArray(block)) return null;
const record = block as Record<string, unknown>;
const tokens = record['tokens'];
const tasks = record['tasks'];
const commits = record['commits'];
if (!isPositiveInt(tokens) || !isPositiveInt(tasks)) return null;
if (typeof commits !== 'number' || !Number.isSafeInteger(commits) || commits < 0) return null;
return { tokens, tasks, commits };
}
/**
* Render an estimate as the YAML block that lands in PLAN.md frontmatter.
* Inverse of parseEstimate over the same value domain — the bijection the
* property test pins.
*/
export function renderEstimate(estimate: PhaseEstimate): string {
const lines = [
'estimate:',
` tokens: ${estimate.tokens}`,
];
if (isPositiveInt(estimate.rawTokens)) lines.push(` raw_tokens: ${estimate.rawTokens}`);
lines.push(` tasks: ${estimate.tasks}`, ` confidence: ${estimate.confidence}`);
return lines.join('\n');
}
/**
* The figure calibration must measure against: the uncalibrated projection when
* the plan recorded one, else the stored value (pre-#2632 plans, where the two
* were the same because no factor had yet been applied).
*/
export function calibrationBasis(estimate: PhaseEstimate): RawTokens {
if (isPositiveInt(estimate.rawTokens)) return estimate.rawTokens;
// THE one legitimate crossover in this module, and the reason asRawTokens()
// refuses a CalibratedTokens rather than being permissive: on a plan written
// before #2632 no factor had been applied yet, so `tokens` IS the raw
// projection. Deliberately an explicit assertion so it stays a single
// auditable line instead of a hole in the brand.
return estimate.tokens as unknown as RawTokens;
}
/** Render an actuals block for SUMMARY.md frontmatter. Inverse of parseActuals. */
export function renderActuals(actuals: PhaseActuals): string {
return [
'actuals:',
` tokens: ${actuals.tokens}`,
` tasks: ${actuals.tasks}`,
` commits: ${actuals.commits}`,
].join('\n');
}
export interface CalibrationDocument {
schema_version: number;
samples: CalibrationSample[];
}
/**
* Parse the persisted calibration document.
*
* This is a trust boundary: the file is on disk, may be hand-edited, and its
* contents steer planning output. Every failure mode degrades to an empty
* sample set rather than throwing or partially trusting — malformed JSON, a
* non-object root, a missing/!== current schema_version, a non-array samples
* field, or individual malformed samples.
*
* A schema_version we do not recognize is refused outright rather than
* best-effort read: a future writer may change the ratio's meaning, and
* misreading it would silently corrupt every subsequent estimate.
*/
export function parseCalibrationDocument(raw: unknown): CalibrationSample[] {
if (typeof raw !== 'string' || raw.trim() === '') return [];
let parsed: unknown;
try {
parsed = JSON.parse(raw);
} catch {
return [];
}
if (parsed === null || typeof parsed !== 'object' || Array.isArray(parsed)) return [];
const doc = parsed as Record<string, unknown>;
if (doc['schema_version'] !== CALIBRATION_SCHEMA_VERSION) return [];
if (!Array.isArray(doc['samples'])) return [];
// Rebuild each sample from its two known fields rather than passing the
// parsed object through — a hostile document cannot smuggle extra keys
// (or a __proto__ payload) into anything downstream.
return (doc['samples'] as unknown[])
.filter(isCalibrationSample)
.map((s) => ({ estimateTokens: s.estimateTokens, actualTokens: s.actualTokens }));
}
/** Serialize a calibration document. Inverse of parseCalibrationDocument. */
export function renderCalibrationDocument(samples: CalibrationSample[]): string {
const doc: CalibrationDocument = { schema_version: CALIBRATION_SCHEMA_VERSION, samples };
return `${JSON.stringify(doc, null, 2)}\n`;
}