14 Commits

Author SHA1 Message Date
Jakub Zych
a9a7a328e6 refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD
across contents and paths, upstream package/repo coordinates -> @golem15/msd-core
and golem15com/msd-core. Deep links into upstream history, sibling upstream
packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is.

Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line,
package/plugin identity, regenerated lockfile, install-tree fixtures, derived
registries and benchmark baseline; migration checksum baseline re-locked
(MSD keeps its own install state, so no install had applied the old sums);
sort-order and regex-escaped expectations in tests adjusted.
2026-10-06 01:47:40 +02:00
Tom Boucher
822934c901 fix(#4794): decision-coverage answers an unmeasured shape on could-not-parse; a non-file context path fails closed (#4889)
* test(#4794): failing-first — could-not-parse must answer an unmeasured shape; a directory context path fails closed

* fix(#4794): could-not-parse answers an unmeasured shape (null counts, unreadable ids, no uncovered); a non-file context path fails closed

* chore(#4794): backfill changeset PR number (4889)

* test(#4794): skip the directory-identity probe when the platform cannot discriminate (windows runner volume collapse, measured)

Two consecutive windows conformance runs failed the probe with measured identical
(dev, ino) for two distinct mkdtemp directories (dev=3606225537, ino=9007199255243448
for both) — a runner-volume property, not a regression in the guard. On such a
platform the guard's identity containment degrades to refuse-everything (fail-closed,
documented); the probe asserts capability, so the honest response is an explicit
t.skip carrying the measurement (ADR-2719 §6), not a red lane for every PR.

---------

Co-authored-by: sim <sim@local>
2026-09-20 02:53:58 -04:00
Tom Boucher
09e1110e68 fix(#4788): code spans in the decision bold lead-in are opaque to the separator grammar (#4883)
* test(#4788): failing-first — code spans in the decision bold lead-in are opaque to the separator grammar

* fix(#4788): the decision lead-in runs are code-span-aware — a backticked span is data, never grammar

* chore(#4788): backfill changeset PR number (4881)

* chore(#4788): correct changeset PR number (4883, was a guessed 4881)

---------

Co-authored-by: sim <sim@local>
2026-09-19 22:46:48 -04:00
Tom Boucher
7bb366e836 fix(#4130): --context flag for check decision-coverage-plan + parseDecisions quadratic-backtracking hardening (#4374)
* test(#4130): failing-first regressions for --context flag + parseDecisions hardening

Block A (flag): check decision-coverage-plan --context <path> must route
identically to the positional form; flag wins over positional context;
valueless --context falls through to the #2770 fail-closed caller error;
verify keeps its positional surface (flag is plan-only). RED on base:
the flag token lands in the args[2] phase slot (false uncovered) or the
args[3] context slot (silent CONTEXT.md-missing skip).

Block B (hardening): regex-lattice asserts pin the atomic-ID wrapper
(?=(X))\1 and the em-dash first-separator narrowing [^*—–]*[—–] plus the
no-adjacent-overlap property; a differential property compares the module
against a frozen copy of the pre-hardening grammars (reference validated
against the base build: 60k generated lines, 0 mismatches); 40k cliff
shapes assert correct outcomes with no wall-time asserts (repo rule).

A12: partitionPredicateArgs keeps one parser behind parsePredicateFlags.

* fix(#4130): --context flag for check decision-coverage-plan + quadratic-backtracking hardening in parseDecisions

(A) check decision-coverage-plan --context <path> — sibling convention
(check predicate, #2008): --flag value pairs parsed by the new shared
partitionPredicateArgs (parsePredicateFlags reimplemented as its flags
half — one parser, cannot diverge), the flag winning over a same-purpose
positional, positionals kept (no sibling deprecates them; the plan-phase
workflow caller passes positionals), valueless --context falls through
to the #2770 fail-closed caller error. Repair of the routing accident
where --context landed in the args[2] phase slot (false uncovered) or
the literal token in the args[3] context slot (silent green skip).

(B) parseDecisions regex seam hardened, byte-identical on all legal
inputs: the three bullet grammars consume the ID atomically via the
(?=(X))\1 lookahead emulation (kills the tail/[^:*]* O(n^2) re-split,
~1.1s @ 40k), and the em-dash first separator narrows [^*]*[—–] to
[^*—–]*[—–] (kills the dash-position O(n^2) retry, ~1.7s @ 40k). Group
indices unchanged (handlers untouched). Pinned by regex-lattice tests,
a differential fast-check property vs the frozen pre-hardening grammars,
and 40k cliff/legal-shape outcome tests (no wall-time asserts per repo
rule — no deterministic engine step counter exists in Node).

* docs+test(#4130): document --context invocation; harden lattice test tooling

- docs/CONFIGURATION.md Decision Coverage Gates: new 'Invoking the plan
  gate directly' block documenting both the positional and --context
  forms, flag precedence, and the valueless-flag fail-closed semantics
  (same place the gate's behavior is documented; sibling check predicate
  documents its flags the same way).
- Two changeset fragments per the maintainer brief (Added: flag; Fixed:
  hardening), PR numbers to be backfilled.
- tests/decisions.test.cjs review fixes: readRegExpTemplate template
  escaping (bare ')' SyntaxError), range-aware lattice checker with
  backreference skip and template unescape, honest A1 contract, lint
  escape warning.

* fix(#4130): valueless --context fails closed per #2770; A8 isolates flag-vs-positional context

Suite-caught fixes from the first verify run:
- cmdDecisionCoveragePlan now refuses a flag-shaped token as the
  positional context path: a bare valueless --context stays a positional
  (sibling parser semantics, unchanged) but reading it as a PATH would
  turn a caller mistake into a silent 'CONTEXT.md missing' green skip —
  exactly what #2770's fail-closed law forbids. Now falls through to
  the missing-context-argument error, as documented.
- A8 test compares decoy-positional+flag against flag-with-phase (phase
  held constant) so the row isolates WHICH context was read; the old
  form compared against a no-phase invocation that could never match.

* chore(#4130): backfill PR number in changeset fragments (PR #4374)

---------

Co-authored-by: sim <sim@local>
2026-09-06 02:55:17 -04:00
Tom Boucher
06eba5fdb0 fix(#4130): parse phase-prefixed decision IDs (D4-01) (#4357)
* test(#4130): failing-first regression for phase-prefixed decision IDs

Add the #4130 matrix: D4-01/D12-01 across all three bullet forms, tags,
discretion, wrapped lead-ins, gate-level plan/verify end-to-end rows, and
parity properties (well-formed digit-prefixed ids parse to their exact id;
a non-digit injected into the prefix fails loud). Update the #2347
non-D-prefix fixture from D5-NN (now a legal grammar) to DEC-NN, and
graduate the representative d5-prefix corpus fixture from could-not-parse
to parsed-but-uncovered.

All new rows are RED against origin/next; they go green with the parser
fix in the next commit.

* fix(#4130): parse phase-prefixed decision IDs (D4-01)

The three declaration grammars, the parse-miss guard, the #3939 join
regexes, and the token evidence all anchored on the literal 'D-' (or
'**D-'), so an ID carrying a digit-run phase prefix between the leading
letter and the hyphen matched nothing — while the #2347 shape detector
correctly called those bullets decision-shaped, collapsing the whole
CONTEXT.md to could-not-parse with 0 extracted instead of a coverage
verdict.

Derive the extractor ID grammar from one shared DECISION_ID_SOURCE
('D[0-9]*-' + the existing alnum tail, full id captured), widen the
guard/join anchors to ID_ATTEMPT_SOURCE (bare 'D-' or a digit-initial
prefix run, so a typo'd 'D4x-01' fails loud while letter-initial prose
like 'Deferred-until' stays none-present), and align the bare-token
evidence. Both gates and the gap-checker share the parser, so all three
surfaces read phase-prefixed decisions now; the gate messages name the
accepted forms including the phase-prefixed one.

* docs(#4130): document the phase-prefixed decision identifier form

The canonical CONTEXT.md reference said decisions carry 'a sequential
D-NN identifier' with no mention of the optional phase-number prefix the
parser now accepts (D4-01) or the alphanumeric tail it always accepted
(D-INFRA-01). Name both in the Decision identifier format section, EN
and ja-JP.

* chore(#4130): changeset

* chore(#4130): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 21:05:19 -04:00
Andreas Brauchli
0ea012c519 fix(#3939): parse decision bullets with a wrapped bold lead-in (#3953)
* fix(#3939): parse decision bullets with a wrapped bold lead-in

parseDecisionLines matched every PHYSICAL line against the three decision-bullet
grammars, and all three require the closing `**` in the same string as the
`- **D-` anchor. A declaration whose bold lead-in wraps across a line break —
the shape discuss-phase itself writes whenever a decision title runs past the
wrap column — matched none of them and fell to the #1365 parse-miss guard, which
forces `could-not-parse` and hard-blocks check.decision-coverage-plan on a
well-formed CONTEXT.md.

Fold physical lines into logical bullets before matching: a declaration whose
bold lead-in is still open at end-of-line absorbs following lines until that run
closes. The three grammars are untouched, so every single-line form parses
exactly as before.

Joining is bounded and preserves the fail-loud contract. A blank or
whitespace-only line, any block-level construct (a list marker of any family,
an ATX heading, a blockquote, a table row), or the end of the block stops it,
and a lead-in that never closes is emitted unchanged — so a genuinely malformed
bullet still reaches the parse-miss guard and still fails loud (#1365), and
cannot be "closed" by an inline `**` belonging to the block below it. The joined
line keeps the first physical line's indent, so the nested cross-reference
signal (#3169) is unchanged. Absorbed lines are scanned once each rather than
re-searching the accumulated candidate, keeping a pathological unterminated run
linear on the plan gate's hot path.

Regression coverage lands in tests/decisions.test.cjs (the owning module's file,
per the regression-test placement policy): all three grammars wrapped, a
three-line wrap, tags/category/continuation preservation, one-line parity
(including inline bold and emphasis inside a wrapped title), CRLF, the
markdown-header path, plus negative proof that every join terminator still
yields could-not-parse and that the FIX-B and #3169 fixtures are unchanged.

Fixes #3939

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3939): add changeset fragment for PR #3953

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3939): property-test the wrap-position invariant

Review follow-up: RULESET.TESTS.property-based-testing requires a parsing /
transformation contract to carry at least one fast-check property asserting a
domain invariant, and the join added by the fix is exactly such a
transformation. The example-based tests pinned four hand-picked wrap points;
these generalize over the whole dimension.

Three properties, on the shared tests/helpers/fast-check-setup.cjs config
(numRuns 200, seeded):

- round-trip: for every grammar (colon-immediate, titled-colon, em-dash), every
  id shape, every tag, with and without a category heading, wrapping the bold
  lead-in at ANY interior space is deepStrictEqual to not wrapping it — where a
  line happens to break carries no information;
- domain invariant: a well-formed wrapped declaration never reaches the
  parse-miss guard (outcome `parsed`) and keeps its declared id;
- fail-loud preservation: an unterminated bold run followed by 0-12 prose lines
  still yields `could-not-parse` with no decision manufactured, however many
  lines the join would have to absorb before giving up.

The corpus is deliberately free of markdown metacharacters: `:` and `*` select a
different grammar (#1639's `[^:*]*` discipline) and a block-construct token
legitimately terminates the join. Both are separate behaviours, example-tested
above; these properties isolate the wrap-position dimension.

Rebuilding the module from `next` with these in place fails 14 (was 12); the two
new failures are the round-trip and never-a-parse-miss properties. The fail-loud
property passes before and after, which is the point of it.

Refs #3939

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3939): fail loud when a wrap splices a decision tag token

Addresses review rounds 2 and 3 on PR #3953.

Folding a soft line break to a single space is markdown's own rule and is
invisible everywhere in a decision bullet except inside the id-adjacent
`[tags]` bracket, which the three grammars turn into `tags` and therefore
into `trackable`. There a spliced space splits one tag token into two
(`[defer` + `red]` -> `defer red`), which does not fail: it parses to a
DIFFERENT tag, silently flipping whether check.decision-coverage-plan
demands coverage for that decision.

The join now stops at such a splice, so the bullet reaches the #1365
parse-miss guard and fails loud instead of guessing. The check is
delimiter-aware, so wraps that land next to `[`, `,` or `]` still join and
still parse identically to the one-line bullet -- a comma-separated tag
list may wrap at any of its separators, across any number of lines. A
bracket further along the title is ordinary text and does not restrict the
join.

Also in this round:

- blockConstructRe's doc comment claimed parity with the sectionizer seam's
  `iterateBullets`, which recognises only the `N. ` ordered form while this
  set also stops at `N) `. The widening is deliberate and one-directional
  (a terminator set may recognise more block openers than a bullet iterator;
  a spare terminator can only make a malformed bullet fail loud, never
  manufacture a decision). Comment corrected to say so, both marker forms
  now tested, and a drift guard asserts the seam still does not yield `N)`
  so the divergence cannot widen silently.
- Documented that the table-row alternative deliberately has no trailing
  whitespace requirement (CommonMark tables may open flush), and that
  over-termination on prose opening `10.` or `|` is accepted fail-loud
  behaviour -- now pinned by a test.
- Coverage the review asked for: a WRAPPED bold lead-in nested under an
  already-open decision (#3169, the existing guard used a single-line nested
  bullet), and title/body whitespace fidelity across every wrap position
  around a double space.
- A fourth fast-check property: wherever a wrap lands inside a `[tags]`
  bracket, the parse either matches the one-line bullet exactly or fails
  loud with nothing extracted -- never a decision whose tags differ.
- Property helpers render through `renderBullet`, which asserts the form
  exists instead of letting an unchecked map lookup yield undefined.

Fail-first: tests/decisions.test.cjs run against origin/next's decisions.cts
fails 18 of 131; against the previous PR head it fails the 2 new tag-splice
guards. All 131 pass with this change. Real-world CONTEXT.md from the report
is unchanged at 37/44 parsed.

Refs #3939

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3939): arm the tag-splice guard on any wrapped line, not just the first

The #3953 round-3 guard read the id-adjacent `[tags]` bracket only from a
bullet's FIRST physical line, via a regex anchored to the bullet start. A
lead-in that wraps twice can open that bracket on a LATER absorbed segment,
where the guard was never armed and `wouldSpliceTagToken` became a no-op:

    - **D-01
      [inform
      ational]: A title.** body text here.

folded to the tag `inform ational` and `trackable: true`, where the one-line
form gives `informational` and `trackable: false` — a silently wrong answer to
the coverage gate, with no thrown error and no parse-miss to signal it. Exactly
the re-classification the round-3 guard exists to prevent, for the case it did
not cover.

`tagBracketOpenAtEolRe` becomes `tagRegionRe`, which asks whether the
id-adjacent bracket REGION is still unsettled rather than whether it opened on
one specific line: group 1 present means the bracket is open, group 1 absent
means the id is read but a `[` may still follow. `joinWrappedBoldLeadIns` keeps
the assembled text in `tagRegion` only while the bracket has yet to open, so a
bracket opening on any segment arms `tagTail`; once armed, the pre-existing
O(1) tail update takes over and `tagRegion` is dropped. A non-empty segment
that is not a bracket-open settles the region immediately, so this bounds the
string to a single extra join and leaves the 5000-line unterminated run linear.

The id class widens to admit an empty id, so a bare `- **D-` still counts as
unsettled. This regex only answers "may an id-adjacent bracket still open
here?", where matching MORE shapes is the conservative direction: an over-broad
match can only make a malformed bullet fail loud, a missed one re-classifies
silently.

The existing property test wraps at exactly one point, and only at spaces —
which round-trip exactly, since the join re-inserts the space it replaced — so
neither the bracket-opens-later state nor an observable splice was reachable
from it. `wrapBoldLeadInMulti` breaks at two or more arbitrary positions after
the id and asserts the same disjunction: parse identically to the one-line
bullet, or fail loud with nothing extracted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BCPSU591zVS9vLPd3gKnqn

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 01:18:36 -04:00
Tom Boucher
6badb839a0 fix(#3514): deny internal fetch hosts; disclose unverified integrity (#3516)
* test(#3514): add failing-first denylist and integrity suites

* fix(#3514): deny internal fetch hosts; disclose unverified integrity

* docs(#3514): trust-model, glossary, and changeset entries

* fix(#3514): scope v6 checks to literals; exact pin kinds in prompt

* chore(#3514): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-14 21:34:29 -04:00
Tom Boucher
fba7c90327 chore(#3484): adr-0174 behavior carry-forward amendment and merge gate (#3507)
* chore(#3484): adr-0174 behavior carry-forward amendment and merge gate

* chore(#3484): regen example context index for new ruleset predicates

* chore(#3484): review fixes - amendment heading per contributor-standards, helper-based fixtures

---------

Co-authored-by: sim <sim@local>
2026-08-14 19:31:48 -04:00
Tom Boucher
470389f3a2 chore(#3212): tokenizer-first for stateful grammars — a shared scanner — Phase 3 (#3424)
* feat(#3414): promote git-cmd.js token-walk into a shared scanner, fix #3169

Phase 3 of epic #3212 (ADR-3212 §4). New src/token-scanner.cts generalizes
hooks/lib/git-cmd.js's proven token-walk (#3129 — "has not re-opened"):
tokenizeShellLike (quote-aware shell tokenizer, byte-identical port) and
indentWidth (bullet-nesting depth).

git-cmd.js migrates onto tokenizeShellLike with zero behavior change
(parity-asserted against every existing #3129 fixture in
tests/worktree-safety.test.cjs's folded block); isGitSubcommand's phases
1-3 (env-prefix skip, executable check, global-option consume) extracted
into skipToSubcommand, shared with the new extractBranchArgument (git
checkout -b / git branch <name>) — a new capability exercising the seam
on the domain the ADR names, not a migration of existing duplicated logic
(none existed).

Fixes #3169: src/decisions.cts's parseDecisionLines couldn't distinguish
a cross-reference bullet nested under an open decision from a fresh
malformed declaration attempt. An earlier bold-run-content-classification
design was tried and disproven against the repo's own existing FIX-B
fixtures (D-02, "no colon no dash") before being adopted — both have
identical shape under any content-only rule. Nesting depth (via
indentWidth) is the actual distinguishing signal: a bullet indented
deeper than the currently-open decision's own bullet is elaboration,
folded into its text like a continuation line, never tested against the
parse-miss guard. A bullet at the same-or-shallower indent is unchanged.

Scope-narrowing disclosed, not silent: of the ADR's four named bugs
(#3197, #3169, #2570, #2528), three no longer need this phase's work.
were independently fixed and closed since the ADR was authored — #2570's
fix is already a correctly-bounded regex per the ADR's own decidability
test (no scanner needed); #2528's fix is a deliberate, twice-reviewed
non-scanner design (its own code comment records a scanner-based attempt
that regressed a symmetric case and was reverted) that this phase does
not disturb. Only #3169 required new work.

get_impact: isGitSubcommand CRITICAL/196 affected symbols,
parseDecisionLines CRITICAL/164 affected symbols (ADR §6 due diligence).

Six-gate ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md,
docs/INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary.

Design: .gsd/phase/chore-3414-tokenizer-first-seam/40-design.md
Test matrix: .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3414): add required fast-check property tests per code review

TESTING-STANDARDS.md:169 requires at least one fast-check property test
for any module that implements parsing — src/token-scanner.cts had none,
an orthogonal Standards-axis review finding. Adds two seeded property
tests (mirroring Phase 1/2's fast-check-setup.cjs convention):
indentWidth counts exactly a generated leading-space run; tokenizeShellLike
round-trips a generated array of whitespace/quote-free words joined with
single spaces.

The design doc's own "no property test needed" rationale was wrong — it
argued no algebraic law applied, but the standard is unconditional for
parsing modules regardless of whether one "feels" applicable. Corrected
in .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md.

Also fixes two Spec-axis wording drifts the same review found between
the design doc and the shipped code (doc-only, no behavior change):
extractBranchArgument's documented signature dropped an unused
subVariants parameter that was never implemented, and the #3169
fail-first fixture description corrected from "15-decision plan via
cmdDecisionCoverageVerify" to the actual compact 3-decision analog via
the real blocking gate, check.decision-coverage-plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): add changeset for #3169 fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): backfill changeset pr number to 3424

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-13 23:08:34 -04:00
Tom Boucher
185da024cb fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block (#2881)
* test(#2770): empty contextPath argument must fail closed, not green-skip the decision-coverage gate

The handler conflated empty-arg (caller error) with file-missing (legitimate skip),
returning passed:true/skipped on an empty argument. Add: empty arg → passed:false;
real-path-to-absent-file → legitimate green skip preserved; omitted arg → fail closed.

* fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block

Handler (check-command-router.cts): split the guard — empty/missing contextPath
argument is a caller error (fail closed, passed:false, mirrors #1365); a real path
whose file genuinely does not exist keeps the legitimate green skip.

Workflow (plan-phase.md): recompute CONTEXT_PATH inside the consuming Bash block
(it was set in the step-1 init block, which does not survive into the separately-
spawned gate block — so the gate ran with an empty arg and silently green-skipped).

* chore(#2770): changeset fragment

* fix(#2770): guard workflow empty-glob case (review blocker) + update drift-guard test

The handler now fails closed on an empty contextPath arg, so the workflow's
unguarded glob (empty when a phase genuinely has no CONTEXT.md) would invoke the
gate with an empty arg → passed:false → exit 1, hard-halting the legitimate
'Continue without context' plan-phase path. Guard the empty-glob case: only run the
gate when a CONTEXT.md actually exists. Update the F1 drift-guard test (which gave
false coverage — it only checked for the ${CONTEXT_PATH} token) to assert the
in-block recompute AND the empty-glob guard.

* fix(#2770): keep plan-phase.md under ADR-857 size cap + ack emitted drift + fix drift-guard window

The workflow fix grew plan-phase.md past the ADR-857 phase-6 size cap (94519B) and
triggered emitted-attribution. Condense adjacent §13a prose/JSON to offset (net
+89B, under cap). Add tests/emitted-drift-ack.json acknowledging the residual growth.
Widen the drift-guard test window (the gate invocation is now nested in the empty-glob
guard, so the old 400-char window missed the glob recompute).

* chore(#2770): backfill changeset PR number (2881)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 16:55:10 -04:00
Tom Boucher
517bae8d6d fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message

Bug: check.decision-coverage-plan's remediation message told the user to
cite decisions "(or body)" but extractPlanDesignatedSections only scanned
<objective>/<tasks>/<task>/<action>. A decision cited in <read_first>,
<behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to
the gate — false BLOCKING coverage gap, plus the message's own fix-hint
sent the user to "the body" where re-citing still failed.

Two-part fix (must change together — that drift was the bug):

1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also
   match <read_first>, <behavior>, <verify>, <acceptance_criteria>,
   <done>. These are all planner-canonical tags the planner is told to
   use (plan-phase.md:830-862, plan-phase.md:772). The body negative-
   lookahead mirrors the opening-tag set so each tag's body is captured
   independently.

2. Correct buildPlanMessage to name ONLY the surfaces the extractor
   actually scans (front-matter must_haves/truths/objective,
   designated markdown headings, and the nine planner-canonical tag
   bodies). The misleading "(or body)" clause is gone.

Also updates the planner's documented contract (agents/gsd-planner.md:69)
and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to
reflect the wider scan.

Regression tests in tests/decisions.test.cjs cover each newly-scanned
tag body, a control (no citation still uncovered), and a message/extractor
parity assertion that names every scanned surface — so the two cannot
drift apart again.

Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage
is a separate command (decision-coverage-verify) checking shipped
artifacts, not plan citations — untouched.

* chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures

gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision-
coverage contract (5 new scanned tag names + heading clarification).
Growth is justified: the contract surface is itself the fix — the prior
text under-described what the gate scans, which was the bug.

Updates:
- tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294)
- 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime)
- tests/fixtures/install-tree/*.json (regenerated by gen:golden)

* fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting

Code review (subagent) flagged a Medium edge-case regression from the
single-alternation regex: when a newly-scanned tag nests inside another
scanned tag, the alternation's negative lookahead halts the outer tag's
body at the inner tag — losing any D-NN citation in the outer tag's
prefix prose. Concretely:

  <action>per D-05 <verify>npm test</verify></action>

  → 3-tag alternation (old):  captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught
  → 9-tag alternation (bug):  captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST

Switches extractXmlTagBodies to per-tag matching: each tag gets its own
regex whose negative-lookahead tempers only against the SAME tag's
reopening. So <verify> inside <action> is absorbed into <action>'s body
(D-05 caught) AND <verify> is matched separately on its own pass.

Per-tag preserves both:
- the reporter's case (sibling tags inside <read_first>)
- nested-tag citations in outer-tag prefix prose
- ReDoS safety (each per-tag regex keeps the #2128 body tempering)

Also adds the reviewer's other requested edge-case tests:
- non-scanned tag (<name>) bearing D-NN must NOT count
- self-closing form <read_first /> safely ignored
- attribute form <verify type="...">D-NN</verify> (canonical planner shape)
- CRLF newlines in tag body do not break capture

* chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md
2026-07-19 23:07:13 -04:00
Tom Boucher
f15eb5f5c9 fix(#2347): make the decision-shape evidence test format-agnostic (#2389)
#1365's fail-loud guard reused the parser's own D- grammar as its evidence test, so a populated <decisions> block using any other ID prefix (e.g. D5-01) was invisible to both parser and guard, collapsing could-not-parse into a clean none-present pass. Add an ID-shaped bold-lead-in probe as format-agnostic evidence on both parse paths; empty/prose scaffolds stay none-present. Graduates the #2371 d5-prefix representative fixture to its expected* assertion.

Closes #2347. Admin-merged (self-review bypass) with full green CI.
2026-07-17 16:25:13 -04:00
Tom Boucher
b205e4c2b2 fix(#1639): parseDecisions handles the titled-colon bullet form (#1665)
* fix(#1639): parseDecisions handles titled-colon bullet form

bulletColonRe anchors on ':**' (colon immediately before close-bold) and bulletEmDashRe
requires an em-dash, so the titled-colon form '- **D-NN: Title.** body' (title between the
colon and the closing **) matched neither and was dropped by the parse-miss guard. When all
decisions used the titled convention, parseDecisions returned 0 and check.decision-coverage-
plan passed vacuously — the same false-coverage failure mode as #1343/#1364/#1365. Add a
third per-form regex bulletTitledColonRe, checked LAST (strict superset of bulletColonRe,
so it only catches bullets the other two miss — minimal blast radius); id + [tags]
trackability honored. Regression folded into decisions.test.cjs: titled-colon parses,
coexists with colon/em-dash, tags, all-titled-13 no longer vacuously 0.

* fix(#1639): tighten titled-colon title to [^:*]* so malformed pre-colon-run bullets still reject

The first cut's title run [^*]* was too permissive: it matched a genuinely-malformed
bullet with a colon in the pre-separator freeform run (e.g. 'D-07 ratio 3:1:**') by
treating the 3:1 colon as the separator, regressing the #1343 parse-miss guard tests.
Tighten the title to [^:*]* (no colon, no star) so the separator colon remains the only
colon permitted before ** — matching bulletColonRe's existing [^:*]* discipline. Valid
titled forms (colon-free titles) still parse; the malformed colon-in-freeform case still
falls through to the parse-miss guard.

* chore(#1639): backfill changeset pr ref to 1665
2026-06-24 13:26:25 -04:00
Tom Boucher
e58b5e1721 fix(#1364): decisions adopt markdown-sectionizer seam + fail-loud coverage gate (epic #1372 T1) (#1386)
* test(#1364,#1365): add decisions regression tests (fail-first proof)

Adds tests/decisions.test.cjs with:
- #1364 recall tests: parseDecisions from markdown-header + em-dash bullets
  (these FAIL on pre-T1 code, proving the bug is present before the fix)
- #1365 fail-loud tests: check.decision-coverage-plan must return passed:false
  for decision-shaped but 0-extracted content (FAIL pre-T1, gate silently passed)
- extractDecisions outcome enum tests (could-not-parse/none-present/parsed)
- Parser QA matrix: CRLF, unicode headings, fenced-code suppression, both bullet forms
- Boundary/threshold tests at limit-1 (0), limit (1)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): adopt markdown-sectionizer seam in decisions.cts; add fail-loud gate

#1364 — Recall: decisions.cts now uses the seam's extractTaggedBlocks and
collectSection for the markdown-header fallback path. Em-dash bullet form
(- **D-NN — title** body) is now recognised alongside the existing colon form.

#1365 — Fail-loud: adds extractDecisions() returning a typed DecisionExtraction
{ decisions, outcome } where outcome is 'parsed' | 'none-present' | 'could-not-parse'.
The blocking gate (cmdDecisionCoveragePlan) now treats could-not-parse as
passed:false with a format-mismatch reason instead of the prior silent passed:true/skip.
gap-checker runGapAnalysis surfaces 'extracted 0 of N — possible format mismatch'
for could-not-parse instead of 'No requirements or decisions to check'.

parseDecisions remains a thin delegate over extractDecisions, so all existing
callers are unaffected.

Seam adoption: stripFencedCode (seam), extractTaggedBlocks(content,'decisions') (seam),
collectSection(content, /decisions?/i, {levelBounded,stripFences}) (seam).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): tighten could-not-parse, parse-miss fail-loud, curly-quote discretion, gap-checker FIX D

FIX A: empty <decisions> scaffolds and all-prose sections no longer return
could-not-parse; outcome is none-present unless the block/section contains
a \bD- token or a parse-miss, preventing false blocks on legitimate phases.

FIX B: parseDecisionLines now tracks parse-misses (D-NN-shaped bullets that
fail both regexes); extractDecisions returns could-not-parse when parseMisses>0
even if some decisions parsed — silent drops no longer mask format errors.

FIX C: curly-quote normalization regex now includes actual U+2018/U+2019
characters so '### Claude's Discretion' (curly apostrophe) correctly yields
trackable:false (regression vs pre-T1 behavior).

FIX D: gap-checker runGapAnalysis surfaces the decision could-not-parse
format-mismatch signal independently of whether requirements items exist —
previously masked inside `if (items.length === 0)`.

Adds 14 behavioral regression tests (fail-first verified manually before fixes).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1365): fail-loud gate on parse-miss regardless of covered decisions

Change the `could-not-parse` guard in `cmdDecisionCoveragePlan` and
`cmdDecisionCoverageVerify` from `decisions.length === 0 && outcome ===
'could-not-parse'` to fire on `outcome === 'could-not-parse'` alone.

Previously a CONTEXT.md with a valid D-01 (covered by the plan) plus a
malformed D-02 (parse-miss) would skip the guard (length === 1), proceed
to coverage, find D-01 covered, and silently return passed:true — hiding
the D-02 parse-miss entirely.

Adds a gate-level fail-first test that places D-01 into a ## Must Haves
section (DESIGNATED_HEADINGS_RE match) so coverage of D-01 would pass on
its own, proving the only path to passed:false is the parse-miss fix.
Also adds the matching verify-side advisory assertion.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1364,#1365): add Fixed changeset (pr:0 placeholder)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1364): backfill changeset PR number (1386)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 12:35:38 -04:00