Commit Graph

5 Commits

Author SHA1 Message Date
Tom Boucher
185da024cb fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block (#2881)
* test(#2770): empty contextPath argument must fail closed, not green-skip the decision-coverage gate

The handler conflated empty-arg (caller error) with file-missing (legitimate skip),
returning passed:true/skipped on an empty argument. Add: empty arg → passed:false;
real-path-to-absent-file → legitimate green skip preserved; omitted arg → fail closed.

* fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block

Handler (check-command-router.cts): split the guard — empty/missing contextPath
argument is a caller error (fail closed, passed:false, mirrors #1365); a real path
whose file genuinely does not exist keeps the legitimate green skip.

Workflow (plan-phase.md): recompute CONTEXT_PATH inside the consuming Bash block
(it was set in the step-1 init block, which does not survive into the separately-
spawned gate block — so the gate ran with an empty arg and silently green-skipped).

* chore(#2770): changeset fragment

* fix(#2770): guard workflow empty-glob case (review blocker) + update drift-guard test

The handler now fails closed on an empty contextPath arg, so the workflow's
unguarded glob (empty when a phase genuinely has no CONTEXT.md) would invoke the
gate with an empty arg → passed:false → exit 1, hard-halting the legitimate
'Continue without context' plan-phase path. Guard the empty-glob case: only run the
gate when a CONTEXT.md actually exists. Update the F1 drift-guard test (which gave
false coverage — it only checked for the ${CONTEXT_PATH} token) to assert the
in-block recompute AND the empty-glob guard.

* fix(#2770): keep plan-phase.md under ADR-857 size cap + ack emitted drift + fix drift-guard window

The workflow fix grew plan-phase.md past the ADR-857 phase-6 size cap (94519B) and
triggered emitted-attribution. Condense adjacent §13a prose/JSON to offset (net
+89B, under cap). Add tests/emitted-drift-ack.json acknowledging the residual growth.
Widen the drift-guard test window (the gate invocation is now nested in the empty-glob
guard, so the old 400-char window missed the glob recompute).

* chore(#2770): backfill changeset PR number (2881)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 16:55:10 -04:00
Tom Boucher
517bae8d6d fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message

Bug: check.decision-coverage-plan's remediation message told the user to
cite decisions "(or body)" but extractPlanDesignatedSections only scanned
<objective>/<tasks>/<task>/<action>. A decision cited in <read_first>,
<behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to
the gate — false BLOCKING coverage gap, plus the message's own fix-hint
sent the user to "the body" where re-citing still failed.

Two-part fix (must change together — that drift was the bug):

1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also
   match <read_first>, <behavior>, <verify>, <acceptance_criteria>,
   <done>. These are all planner-canonical tags the planner is told to
   use (plan-phase.md:830-862, plan-phase.md:772). The body negative-
   lookahead mirrors the opening-tag set so each tag's body is captured
   independently.

2. Correct buildPlanMessage to name ONLY the surfaces the extractor
   actually scans (front-matter must_haves/truths/objective,
   designated markdown headings, and the nine planner-canonical tag
   bodies). The misleading "(or body)" clause is gone.

Also updates the planner's documented contract (agents/gsd-planner.md:69)
and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to
reflect the wider scan.

Regression tests in tests/decisions.test.cjs cover each newly-scanned
tag body, a control (no citation still uncovered), and a message/extractor
parity assertion that names every scanned surface — so the two cannot
drift apart again.

Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage
is a separate command (decision-coverage-verify) checking shipped
artifacts, not plan citations — untouched.

* chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures

gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision-
coverage contract (5 new scanned tag names + heading clarification).
Growth is justified: the contract surface is itself the fix — the prior
text under-described what the gate scans, which was the bug.

Updates:
- tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294)
- 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime)
- tests/fixtures/install-tree/*.json (regenerated by gen:golden)

* fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting

Code review (subagent) flagged a Medium edge-case regression from the
single-alternation regex: when a newly-scanned tag nests inside another
scanned tag, the alternation's negative lookahead halts the outer tag's
body at the inner tag — losing any D-NN citation in the outer tag's
prefix prose. Concretely:

  <action>per D-05 <verify>npm test</verify></action>

  → 3-tag alternation (old):  captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught
  → 9-tag alternation (bug):  captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST

Switches extractXmlTagBodies to per-tag matching: each tag gets its own
regex whose negative-lookahead tempers only against the SAME tag's
reopening. So <verify> inside <action> is absorbed into <action>'s body
(D-05 caught) AND <verify> is matched separately on its own pass.

Per-tag preserves both:
- the reporter's case (sibling tags inside <read_first>)
- nested-tag citations in outer-tag prefix prose
- ReDoS safety (each per-tag regex keeps the #2128 body tempering)

Also adds the reviewer's other requested edge-case tests:
- non-scanned tag (<name>) bearing D-NN must NOT count
- self-closing form <read_first /> safely ignored
- attribute form <verify type="...">D-NN</verify> (canonical planner shape)
- CRLF newlines in tag body do not break capture

* chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md
2026-07-19 23:07:13 -04:00
Tom Boucher
f15eb5f5c9 fix(#2347): make the decision-shape evidence test format-agnostic (#2389)
#1365's fail-loud guard reused the parser's own D- grammar as its evidence test, so a populated <decisions> block using any other ID prefix (e.g. D5-01) was invisible to both parser and guard, collapsing could-not-parse into a clean none-present pass. Add an ID-shaped bold-lead-in probe as format-agnostic evidence on both parse paths; empty/prose scaffolds stay none-present. Graduates the #2371 d5-prefix representative fixture to its expected* assertion.

Closes #2347. Admin-merged (self-review bypass) with full green CI.
2026-07-17 16:25:13 -04:00
Tom Boucher
b205e4c2b2 fix(#1639): parseDecisions handles the titled-colon bullet form (#1665)
* fix(#1639): parseDecisions handles titled-colon bullet form

bulletColonRe anchors on ':**' (colon immediately before close-bold) and bulletEmDashRe
requires an em-dash, so the titled-colon form '- **D-NN: Title.** body' (title between the
colon and the closing **) matched neither and was dropped by the parse-miss guard. When all
decisions used the titled convention, parseDecisions returned 0 and check.decision-coverage-
plan passed vacuously — the same false-coverage failure mode as #1343/#1364/#1365. Add a
third per-form regex bulletTitledColonRe, checked LAST (strict superset of bulletColonRe,
so it only catches bullets the other two miss — minimal blast radius); id + [tags]
trackability honored. Regression folded into decisions.test.cjs: titled-colon parses,
coexists with colon/em-dash, tags, all-titled-13 no longer vacuously 0.

* fix(#1639): tighten titled-colon title to [^:*]* so malformed pre-colon-run bullets still reject

The first cut's title run [^*]* was too permissive: it matched a genuinely-malformed
bullet with a colon in the pre-separator freeform run (e.g. 'D-07 ratio 3:1:**') by
treating the 3:1 colon as the separator, regressing the #1343 parse-miss guard tests.
Tighten the title to [^:*]* (no colon, no star) so the separator colon remains the only
colon permitted before ** — matching bulletColonRe's existing [^:*]* discipline. Valid
titled forms (colon-free titles) still parse; the malformed colon-in-freeform case still
falls through to the parse-miss guard.

* chore(#1639): backfill changeset pr ref to 1665
2026-06-24 13:26:25 -04:00
Tom Boucher
e58b5e1721 fix(#1364): decisions adopt markdown-sectionizer seam + fail-loud coverage gate (epic #1372 T1) (#1386)
* test(#1364,#1365): add decisions regression tests (fail-first proof)

Adds tests/decisions.test.cjs with:
- #1364 recall tests: parseDecisions from markdown-header + em-dash bullets
  (these FAIL on pre-T1 code, proving the bug is present before the fix)
- #1365 fail-loud tests: check.decision-coverage-plan must return passed:false
  for decision-shaped but 0-extracted content (FAIL pre-T1, gate silently passed)
- extractDecisions outcome enum tests (could-not-parse/none-present/parsed)
- Parser QA matrix: CRLF, unicode headings, fenced-code suppression, both bullet forms
- Boundary/threshold tests at limit-1 (0), limit (1)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): adopt markdown-sectionizer seam in decisions.cts; add fail-loud gate

#1364 — Recall: decisions.cts now uses the seam's extractTaggedBlocks and
collectSection for the markdown-header fallback path. Em-dash bullet form
(- **D-NN — title** body) is now recognised alongside the existing colon form.

#1365 — Fail-loud: adds extractDecisions() returning a typed DecisionExtraction
{ decisions, outcome } where outcome is 'parsed' | 'none-present' | 'could-not-parse'.
The blocking gate (cmdDecisionCoveragePlan) now treats could-not-parse as
passed:false with a format-mismatch reason instead of the prior silent passed:true/skip.
gap-checker runGapAnalysis surfaces 'extracted 0 of N — possible format mismatch'
for could-not-parse instead of 'No requirements or decisions to check'.

parseDecisions remains a thin delegate over extractDecisions, so all existing
callers are unaffected.

Seam adoption: stripFencedCode (seam), extractTaggedBlocks(content,'decisions') (seam),
collectSection(content, /decisions?/i, {levelBounded,stripFences}) (seam).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): tighten could-not-parse, parse-miss fail-loud, curly-quote discretion, gap-checker FIX D

FIX A: empty <decisions> scaffolds and all-prose sections no longer return
could-not-parse; outcome is none-present unless the block/section contains
a \bD- token or a parse-miss, preventing false blocks on legitimate phases.

FIX B: parseDecisionLines now tracks parse-misses (D-NN-shaped bullets that
fail both regexes); extractDecisions returns could-not-parse when parseMisses>0
even if some decisions parsed — silent drops no longer mask format errors.

FIX C: curly-quote normalization regex now includes actual U+2018/U+2019
characters so '### Claude's Discretion' (curly apostrophe) correctly yields
trackable:false (regression vs pre-T1 behavior).

FIX D: gap-checker runGapAnalysis surfaces the decision could-not-parse
format-mismatch signal independently of whether requirements items exist —
previously masked inside `if (items.length === 0)`.

Adds 14 behavioral regression tests (fail-first verified manually before fixes).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1365): fail-loud gate on parse-miss regardless of covered decisions

Change the `could-not-parse` guard in `cmdDecisionCoveragePlan` and
`cmdDecisionCoverageVerify` from `decisions.length === 0 && outcome ===
'could-not-parse'` to fire on `outcome === 'could-not-parse'` alone.

Previously a CONTEXT.md with a valid D-01 (covered by the plan) plus a
malformed D-02 (parse-miss) would skip the guard (length === 1), proceed
to coverage, find D-01 covered, and silently return passed:true — hiding
the D-02 parse-miss entirely.

Adds a gate-level fail-first test that places D-01 into a ## Must Haves
section (DESIGNATED_HEADINGS_RE match) so coverage of D-01 would pass on
its own, proving the only path to passed:false is the parse-miss fix.
Also adds the matching verify-side advisory assertion.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1364,#1365): add Fixed changeset (pr:0 placeholder)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1364): backfill changeset PR number (1386)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 12:35:38 -04:00