* test(#3882): failing-first rows for sentinel phases skewing calibration
Adds A1a/A1b/A2/A3 to tests/estimate-calibrate.test.cjs, the module's
existing test file, rather than a new bug-NNNN file. collectCalibrationSamples
(src/estimate-cli.cts:206) does a raw readdirSync over .planning/phases and
never applies isSentinelPhaseId, so a sentinel phase (milestone 0 or 999)
carrying a PLAN estimate / SUMMARY actuals pair contributes a phantom
calibration sample.
computeCalibration is median-based, so a single 50x outlier among three
samples leaves the factor unmoved — asserting "the factor is unchanged"
against one sentinel would pass on the broken code for the wrong reason.
Each row instead asserts the WHOLE computed CalibrationResult object
(factor, applied, confidence, sampleCount, clamped) for a sentinel-free
project against its sentinel-injected twin:
- A1a: one sentinel flips applied false->true and confidence low->med on
phantom evidence (calibration switches on with zero real signal).
- A1b: two sentinels corrupt the factor itself (1 -> 3, clamped false->true).
- A2: the sentinel's own sample is verified absent from the returned list.
- A3: the two genuine phases still contribute their own unchanged samples
(regression pin — stops A1/A2 passing by filtering everything).
Verified RED on today's code (node tests/estimate-calibrate.test.cjs):
A1a/A1b/A2 fail with the exact differing objects; A3 and all pre-existing
rows in the file remain green (no collateral).
Refs #3882
* feat(#3882): route phase enumeration through its owner and name the sentinel axis
Task 1: collectCalibrationSamples (src/estimate-cli.cts) hand-rolled a raw readdirSync over .planning/phases, treating every directory (including sentinel phases, milestone 0/999) as a completed phase and feeding phantom PLAN/SUMMARY samples into the estimation calibration factor. Routed through the existing owner, listMilestonePhaseDirs(phasesRoot) with no cwd -- already 'all milestones, sentinels excluded', exactly the combination this caller needs; no new API was required for this half. It now also surfaces the scope discriminator: an unreadable phases directory throws PhasesUnreadableError instead of silently returning zero samples, and cmdEstimateCalibrate reports it via a new ERROR_REASON.ESTIMATE_PHASES_UNREADABLE instead of persisting a phantom empty calibration document.
Task 2: added listAllPhaseDirs(phasesDir, { includeSentinels }) to src/phase-locator.cts -- the one genuinely missing axis: 'physical set, sentinels INCLUDED'. includeSentinels has no default and is required, so a call site cannot obtain sentinel-inclusion by omission (compile-time refusal, not just documentation). Mirrors listMilestonePhaseDirs's absent/unreadable scope handling.
Task 3: migrated the two exemptions whose written reason maps cleanly onto 'physical set, sentinels included' -- cmdRoadmapAnalyze's _phaseDirNames (src/roadmap.cts) and cmdInitMilestoneOp's diskPhaseDirs (src/init.cts), both heading->directory lookup indexes. Left the rest: archivePhaseDirectories's own body has no readdirSync to migrate (its callers already resolve dirs before calling it, and both current callers deliberately EXCLUDE sentinels -- migrating it would be an unauthorized behavior change, not an API swap); cmdValidateHealth's exemption is vestigial (its actual physical-set sweep already lives in planning-snapshot.cts's buildAllPhaseDirNamesField, a pre-existing near-duplicate of the new axis, flagged as a finding, not restructured); cmdPhasesClear/cmdMilestoneComplete/cmdVerifySchemaDrift/detectHasPriorPhases/detectUiPhaseActive want a different combination (sentinels excluded, or a single-phase lookup) and are unaffected.
Task 4: detector 2 (sentinel literal) is untouched and retained. Removed exemption entries only for the two migrated call sites; every other function-scoped exemption is preserved. Guard exits 0.
Refs #3882
* refactor(#3882): delegate the snapshot phase-dir scan to its owner
buildAllPhaseDirNamesField duplicated listAllPhaseDirs's own
readdirSync + directory-filter + absent/unreadable handling — the
'one implementation per rule' defect ADR-3473 SS8.3 names, introduced
by this branch's own #3882 work. Delegate to listAllPhaseDirs and
re-apply the field's existing lexicographic sort on top, since W007's
observable order must not change.
Refs #3882
* docs(#3882): document the sentinel axis and the enumeration consolidation
Records listAllPhaseDirs in the Phase Locator glossary entry, and the fact
that the owner already answers the all-milestones sentinel-free question when
called without a cwd -- the call collectCalibrationSamples was missing.
Also notes that buildAllPhaseDirNamesField now delegates rather than carrying a
second readdir, and that exactly one readdirSync over the phases directory
remains across the two modules.
Refs #3882
* test(#3882): close review findings — real order proof, unreadable coverage, collision fixtures
Refs #3882
* chore(#3882): backfill changeset PR number
Refs #3882
---------
Co-authored-by: sim <sim@local>
* test(#2671): add failing-first brand-typing compile fixtures
* feat(#2671): brand raw vs calibrated token types
* refactor(#2671): hoist type-compile into a before() hook
Two review responses:
- The fixture compile ran in the describe() body, so it executed at
collection time even when the block was filtered out, and a failed
precondition collapsed eight independent assertions into one opaque
describe-level failure. A before() hook is this repo's documented
idiom and preserves per-test granularity.
- parseTokensFlag now records WHY it returns an unbranded number: it
validates the magnitude of --tokens, but the basis is decided by
--calibrated, so branding here would be wrong for half its callers.
The assertion belongs to cmdEstimateCheck, its only caller.
* test(#2671): pin each brand diagnostic to its OFFENDING marker
Adversarial review demonstrated that asserting only exactly-one-diagnostic-
at-code-N is not airtight. Repairing a fixture's brand violation while
injecting an unrelated error of the same code (a string passed as the
budget argument) still yielded exactly one TS2345, so the fixture would
have reported green while no longer testing its regression at all.
Each bad-* fixture now routes its violating value through a const named
OFFENDING, and the test asserts the diagnostic's start offset falls inside
that node — located through the AST, so it survives reformatting and never
pattern-matches source text. Replaying the proof-of-concept against the new
assertion rejects it: the diagnostic lands on the budget literal, not the
marker.
Also corrects a doc comment that claimed the program type-checks all of
src/; it covers phase-estimation.cts and its transitive dependencies.
* chore(#2671): backfill changeset PR number (#2676)
* feat(#2632): record executor actuals and close the estimate calibration loop
* fix(#2632): calibrate against the raw projection so the loop converges
* test(#2632): add closed-loop convergence guard and codify the feedback-loop rule
* fix(#2632): pair calibration samples per plan; atomic write; amend adr
* chore(#2632): backfill changeset pr to 2672
* fix(#2632): retry renameSync on transient windows errnos and clean up the temp
* test(#2631): failing-first planner estimate emission and over-budget surfacing
* feat(#2631): emit plan estimate and surface the over-budget split recommendation
* fix(#2631): extract sizing prose to references to fit planner and plan-phase caps
* fix(#2631): move estimate check to plan-checker; fix template regex and caps
* fix(#2631): restore ALWAYS split literal and keep gsd_run after the launcher preamble
* fix(#2631): invoke estimate-check after the launcher preamble in plan-checker
* fix(#2631): stop double-applying calibration; repair COMMANDS table and stale reference
* chore(#2631): backfill changeset pr to 2670
* chore(#2631): backfill changeset pr to 2670