* test(#2671): add failing-first brand-typing compile fixtures
* feat(#2671): brand raw vs calibrated token types
* refactor(#2671): hoist type-compile into a before() hook
Two review responses:
- The fixture compile ran in the describe() body, so it executed at
collection time even when the block was filtered out, and a failed
precondition collapsed eight independent assertions into one opaque
describe-level failure. A before() hook is this repo's documented
idiom and preserves per-test granularity.
- parseTokensFlag now records WHY it returns an unbranded number: it
validates the magnitude of --tokens, but the basis is decided by
--calibrated, so branding here would be wrong for half its callers.
The assertion belongs to cmdEstimateCheck, its only caller.
* test(#2671): pin each brand diagnostic to its OFFENDING marker
Adversarial review demonstrated that asserting only exactly-one-diagnostic-
at-code-N is not airtight. Repairing a fixture's brand violation while
injecting an unrelated error of the same code (a string passed as the
budget argument) still yielded exactly one TS2345, so the fixture would
have reported green while no longer testing its regression at all.
Each bad-* fixture now routes its violating value through a const named
OFFENDING, and the test asserts the diagnostic's start offset falls inside
that node — located through the AST, so it survives reformatting and never
pattern-matches source text. Replaying the proof-of-concept against the new
assertion rejects it: the diagnostic lands on the budget literal, not the
marker.
Also corrects a doc comment that claimed the program type-checks all of
src/; it covers phase-estimation.cts and its transitive dependencies.
* chore(#2671): backfill changeset PR number (#2676)