next
14 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a9a7a328e6 |
refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted. |
||
|
|
822934c901 |
fix(#4794): decision-coverage answers an unmeasured shape on could-not-parse; a non-file context path fails closed (#4889)
* test(#4794): failing-first — could-not-parse must answer an unmeasured shape; a directory context path fails closed * fix(#4794): could-not-parse answers an unmeasured shape (null counts, unreadable ids, no uncovered); a non-file context path fails closed * chore(#4794): backfill changeset PR number (4889) * test(#4794): skip the directory-identity probe when the platform cannot discriminate (windows runner volume collapse, measured) Two consecutive windows conformance runs failed the probe with measured identical (dev, ino) for two distinct mkdtemp directories (dev=3606225537, ino=9007199255243448 for both) — a runner-volume property, not a regression in the guard. On such a platform the guard's identity containment degrades to refuse-everything (fail-closed, documented); the probe asserts capability, so the honest response is an explicit t.skip carrying the measurement (ADR-2719 §6), not a red lane for every PR. --------- Co-authored-by: sim <sim@local> |
||
|
|
09e1110e68 |
fix(#4788): code spans in the decision bold lead-in are opaque to the separator grammar (#4883)
* test(#4788): failing-first — code spans in the decision bold lead-in are opaque to the separator grammar * fix(#4788): the decision lead-in runs are code-span-aware — a backticked span is data, never grammar * chore(#4788): backfill changeset PR number (4881) * chore(#4788): correct changeset PR number (4883, was a guessed 4881) --------- Co-authored-by: sim <sim@local> |
||
|
|
7bb366e836 |
fix(#4130): --context flag for check decision-coverage-plan + parseDecisions quadratic-backtracking hardening (#4374)
* test(#4130): failing-first regressions for --context flag + parseDecisions hardening Block A (flag): check decision-coverage-plan --context <path> must route identically to the positional form; flag wins over positional context; valueless --context falls through to the #2770 fail-closed caller error; verify keeps its positional surface (flag is plan-only). RED on base: the flag token lands in the args[2] phase slot (false uncovered) or the args[3] context slot (silent CONTEXT.md-missing skip). Block B (hardening): regex-lattice asserts pin the atomic-ID wrapper (?=(X))\1 and the em-dash first-separator narrowing [^*—–]*[—–] plus the no-adjacent-overlap property; a differential property compares the module against a frozen copy of the pre-hardening grammars (reference validated against the base build: 60k generated lines, 0 mismatches); 40k cliff shapes assert correct outcomes with no wall-time asserts (repo rule). A12: partitionPredicateArgs keeps one parser behind parsePredicateFlags. * fix(#4130): --context flag for check decision-coverage-plan + quadratic-backtracking hardening in parseDecisions (A) check decision-coverage-plan --context <path> — sibling convention (check predicate, #2008): --flag value pairs parsed by the new shared partitionPredicateArgs (parsePredicateFlags reimplemented as its flags half — one parser, cannot diverge), the flag winning over a same-purpose positional, positionals kept (no sibling deprecates them; the plan-phase workflow caller passes positionals), valueless --context falls through to the #2770 fail-closed caller error. Repair of the routing accident where --context landed in the args[2] phase slot (false uncovered) or the literal token in the args[3] context slot (silent green skip). (B) parseDecisions regex seam hardened, byte-identical on all legal inputs: the three bullet grammars consume the ID atomically via the (?=(X))\1 lookahead emulation (kills the tail/[^:*]* O(n^2) re-split, ~1.1s @ 40k), and the em-dash first separator narrows [^*]*[—–] to [^*—–]*[—–] (kills the dash-position O(n^2) retry, ~1.7s @ 40k). Group indices unchanged (handlers untouched). Pinned by regex-lattice tests, a differential fast-check property vs the frozen pre-hardening grammars, and 40k cliff/legal-shape outcome tests (no wall-time asserts per repo rule — no deterministic engine step counter exists in Node). * docs+test(#4130): document --context invocation; harden lattice test tooling - docs/CONFIGURATION.md Decision Coverage Gates: new 'Invoking the plan gate directly' block documenting both the positional and --context forms, flag precedence, and the valueless-flag fail-closed semantics (same place the gate's behavior is documented; sibling check predicate documents its flags the same way). - Two changeset fragments per the maintainer brief (Added: flag; Fixed: hardening), PR numbers to be backfilled. - tests/decisions.test.cjs review fixes: readRegExpTemplate template escaping (bare ')' SyntaxError), range-aware lattice checker with backreference skip and template unescape, honest A1 contract, lint escape warning. * fix(#4130): valueless --context fails closed per #2770; A8 isolates flag-vs-positional context Suite-caught fixes from the first verify run: - cmdDecisionCoveragePlan now refuses a flag-shaped token as the positional context path: a bare valueless --context stays a positional (sibling parser semantics, unchanged) but reading it as a PATH would turn a caller mistake into a silent 'CONTEXT.md missing' green skip — exactly what #2770's fail-closed law forbids. Now falls through to the missing-context-argument error, as documented. - A8 test compares decoy-positional+flag against flag-with-phase (phase held constant) so the row isolates WHICH context was read; the old form compared against a no-phase invocation that could never match. * chore(#4130): backfill PR number in changeset fragments (PR #4374) --------- Co-authored-by: sim <sim@local> |
||
|
|
06eba5fdb0 |
fix(#4130): parse phase-prefixed decision IDs (D4-01) (#4357)
* test(#4130): failing-first regression for phase-prefixed decision IDs Add the #4130 matrix: D4-01/D12-01 across all three bullet forms, tags, discretion, wrapped lead-ins, gate-level plan/verify end-to-end rows, and parity properties (well-formed digit-prefixed ids parse to their exact id; a non-digit injected into the prefix fails loud). Update the #2347 non-D-prefix fixture from D5-NN (now a legal grammar) to DEC-NN, and graduate the representative d5-prefix corpus fixture from could-not-parse to parsed-but-uncovered. All new rows are RED against origin/next; they go green with the parser fix in the next commit. * fix(#4130): parse phase-prefixed decision IDs (D4-01) The three declaration grammars, the parse-miss guard, the #3939 join regexes, and the token evidence all anchored on the literal 'D-' (or '**D-'), so an ID carrying a digit-run phase prefix between the leading letter and the hyphen matched nothing — while the #2347 shape detector correctly called those bullets decision-shaped, collapsing the whole CONTEXT.md to could-not-parse with 0 extracted instead of a coverage verdict. Derive the extractor ID grammar from one shared DECISION_ID_SOURCE ('D[0-9]*-' + the existing alnum tail, full id captured), widen the guard/join anchors to ID_ATTEMPT_SOURCE (bare 'D-' or a digit-initial prefix run, so a typo'd 'D4x-01' fails loud while letter-initial prose like 'Deferred-until' stays none-present), and align the bare-token evidence. Both gates and the gap-checker share the parser, so all three surfaces read phase-prefixed decisions now; the gate messages name the accepted forms including the phase-prefixed one. * docs(#4130): document the phase-prefixed decision identifier form The canonical CONTEXT.md reference said decisions carry 'a sequential D-NN identifier' with no mention of the optional phase-number prefix the parser now accepts (D4-01) or the alphanumeric tail it always accepted (D-INFRA-01). Name both in the Decision identifier format section, EN and ja-JP. * chore(#4130): changeset * chore(#4130): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
0ea012c519 |
fix(#3939): parse decision bullets with a wrapped bold lead-in (#3953)
* fix(#3939): parse decision bullets with a wrapped bold lead-in parseDecisionLines matched every PHYSICAL line against the three decision-bullet grammars, and all three require the closing `**` in the same string as the `- **D-` anchor. A declaration whose bold lead-in wraps across a line break — the shape discuss-phase itself writes whenever a decision title runs past the wrap column — matched none of them and fell to the #1365 parse-miss guard, which forces `could-not-parse` and hard-blocks check.decision-coverage-plan on a well-formed CONTEXT.md. Fold physical lines into logical bullets before matching: a declaration whose bold lead-in is still open at end-of-line absorbs following lines until that run closes. The three grammars are untouched, so every single-line form parses exactly as before. Joining is bounded and preserves the fail-loud contract. A blank or whitespace-only line, any block-level construct (a list marker of any family, an ATX heading, a blockquote, a table row), or the end of the block stops it, and a lead-in that never closes is emitted unchanged — so a genuinely malformed bullet still reaches the parse-miss guard and still fails loud (#1365), and cannot be "closed" by an inline `**` belonging to the block below it. The joined line keeps the first physical line's indent, so the nested cross-reference signal (#3169) is unchanged. Absorbed lines are scanned once each rather than re-searching the accumulated candidate, keeping a pathological unterminated run linear on the plan gate's hot path. Regression coverage lands in tests/decisions.test.cjs (the owning module's file, per the regression-test placement policy): all three grammars wrapped, a three-line wrap, tags/category/continuation preservation, one-line parity (including inline bold and emphasis inside a wrapped title), CRLF, the markdown-header path, plus negative proof that every join terminator still yields could-not-parse and that the FIX-B and #3169 fixtures are unchanged. Fixes #3939 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3939): add changeset fragment for PR #3953 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3939): property-test the wrap-position invariant Review follow-up: RULESET.TESTS.property-based-testing requires a parsing / transformation contract to carry at least one fast-check property asserting a domain invariant, and the join added by the fix is exactly such a transformation. The example-based tests pinned four hand-picked wrap points; these generalize over the whole dimension. Three properties, on the shared tests/helpers/fast-check-setup.cjs config (numRuns 200, seeded): - round-trip: for every grammar (colon-immediate, titled-colon, em-dash), every id shape, every tag, with and without a category heading, wrapping the bold lead-in at ANY interior space is deepStrictEqual to not wrapping it — where a line happens to break carries no information; - domain invariant: a well-formed wrapped declaration never reaches the parse-miss guard (outcome `parsed`) and keeps its declared id; - fail-loud preservation: an unterminated bold run followed by 0-12 prose lines still yields `could-not-parse` with no decision manufactured, however many lines the join would have to absorb before giving up. The corpus is deliberately free of markdown metacharacters: `:` and `*` select a different grammar (#1639's `[^:*]*` discipline) and a block-construct token legitimately terminates the join. Both are separate behaviours, example-tested above; these properties isolate the wrap-position dimension. Rebuilding the module from `next` with these in place fails 14 (was 12); the two new failures are the round-trip and never-a-parse-miss properties. The fail-loud property passes before and after, which is the point of it. Refs #3939 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3939): fail loud when a wrap splices a decision tag token Addresses review rounds 2 and 3 on PR #3953. Folding a soft line break to a single space is markdown's own rule and is invisible everywhere in a decision bullet except inside the id-adjacent `[tags]` bracket, which the three grammars turn into `tags` and therefore into `trackable`. There a spliced space splits one tag token into two (`[defer` + `red]` -> `defer red`), which does not fail: it parses to a DIFFERENT tag, silently flipping whether check.decision-coverage-plan demands coverage for that decision. The join now stops at such a splice, so the bullet reaches the #1365 parse-miss guard and fails loud instead of guessing. The check is delimiter-aware, so wraps that land next to `[`, `,` or `]` still join and still parse identically to the one-line bullet -- a comma-separated tag list may wrap at any of its separators, across any number of lines. A bracket further along the title is ordinary text and does not restrict the join. Also in this round: - blockConstructRe's doc comment claimed parity with the sectionizer seam's `iterateBullets`, which recognises only the `N. ` ordered form while this set also stops at `N) `. The widening is deliberate and one-directional (a terminator set may recognise more block openers than a bullet iterator; a spare terminator can only make a malformed bullet fail loud, never manufacture a decision). Comment corrected to say so, both marker forms now tested, and a drift guard asserts the seam still does not yield `N)` so the divergence cannot widen silently. - Documented that the table-row alternative deliberately has no trailing whitespace requirement (CommonMark tables may open flush), and that over-termination on prose opening `10.` or `|` is accepted fail-loud behaviour -- now pinned by a test. - Coverage the review asked for: a WRAPPED bold lead-in nested under an already-open decision (#3169, the existing guard used a single-line nested bullet), and title/body whitespace fidelity across every wrap position around a double space. - A fourth fast-check property: wherever a wrap lands inside a `[tags]` bracket, the parse either matches the one-line bullet exactly or fails loud with nothing extracted -- never a decision whose tags differ. - Property helpers render through `renderBullet`, which asserts the form exists instead of letting an unchecked map lookup yield undefined. Fail-first: tests/decisions.test.cjs run against origin/next's decisions.cts fails 18 of 131; against the previous PR head it fails the 2 new tag-splice guards. All 131 pass with this change. Real-world CONTEXT.md from the report is unchanged at 37/44 parsed. Refs #3939 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3939): arm the tag-splice guard on any wrapped line, not just the first The #3953 round-3 guard read the id-adjacent `[tags]` bracket only from a bullet's FIRST physical line, via a regex anchored to the bullet start. A lead-in that wraps twice can open that bracket on a LATER absorbed segment, where the guard was never armed and `wouldSpliceTagToken` became a no-op: - **D-01 [inform ational]: A title.** body text here. folded to the tag `inform ational` and `trackable: true`, where the one-line form gives `informational` and `trackable: false` — a silently wrong answer to the coverage gate, with no thrown error and no parse-miss to signal it. Exactly the re-classification the round-3 guard exists to prevent, for the case it did not cover. `tagBracketOpenAtEolRe` becomes `tagRegionRe`, which asks whether the id-adjacent bracket REGION is still unsettled rather than whether it opened on one specific line: group 1 present means the bracket is open, group 1 absent means the id is read but a `[` may still follow. `joinWrappedBoldLeadIns` keeps the assembled text in `tagRegion` only while the bracket has yet to open, so a bracket opening on any segment arms `tagTail`; once armed, the pre-existing O(1) tail update takes over and `tagRegion` is dropped. A non-empty segment that is not a bracket-open settles the region immediately, so this bounds the string to a single extra join and leaves the 5000-line unterminated run linear. The id class widens to admit an empty id, so a bare `- **D-` still counts as unsettled. This regex only answers "may an id-adjacent bracket still open here?", where matching MORE shapes is the conservative direction: an over-broad match can only make a malformed bullet fail loud, a missed one re-classifies silently. The existing property test wraps at exactly one point, and only at spaces — which round-trip exactly, since the join re-inserts the space it replaced — so neither the bracket-opens-later state nor an observable splice was reachable from it. `wrapBoldLeadInMulti` breaks at two or more arbitrary positions after the id and asserts the same disjunction: parse identically to the one-line bullet, or fail loud with nothing extracted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BCPSU591zVS9vLPd3gKnqn --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
6badb839a0 |
fix(#3514): deny internal fetch hosts; disclose unverified integrity (#3516)
* test(#3514): add failing-first denylist and integrity suites * fix(#3514): deny internal fetch hosts; disclose unverified integrity * docs(#3514): trust-model, glossary, and changeset entries * fix(#3514): scope v6 checks to literals; exact pin kinds in prompt * chore(#3514): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
fba7c90327 |
chore(#3484): adr-0174 behavior carry-forward amendment and merge gate (#3507)
* chore(#3484): adr-0174 behavior carry-forward amendment and merge gate * chore(#3484): regen example context index for new ruleset predicates * chore(#3484): review fixes - amendment heading per contributor-standards, helper-based fixtures --------- Co-authored-by: sim <sim@local> |
||
|
|
470389f3a2 |
chore(#3212): tokenizer-first for stateful grammars — a shared scanner — Phase 3 (#3424)
* feat(#3414): promote git-cmd.js token-walk into a shared scanner, fix #3169 Phase 3 of epic #3212 (ADR-3212 §4). New src/token-scanner.cts generalizes hooks/lib/git-cmd.js's proven token-walk (#3129 — "has not re-opened"): tokenizeShellLike (quote-aware shell tokenizer, byte-identical port) and indentWidth (bullet-nesting depth). git-cmd.js migrates onto tokenizeShellLike with zero behavior change (parity-asserted against every existing #3129 fixture in tests/worktree-safety.test.cjs's folded block); isGitSubcommand's phases 1-3 (env-prefix skip, executable check, global-option consume) extracted into skipToSubcommand, shared with the new extractBranchArgument (git checkout -b / git branch <name>) — a new capability exercising the seam on the domain the ADR names, not a migration of existing duplicated logic (none existed). Fixes #3169: src/decisions.cts's parseDecisionLines couldn't distinguish a cross-reference bullet nested under an open decision from a fresh malformed declaration attempt. An earlier bold-run-content-classification design was tried and disproven against the repo's own existing FIX-B fixtures (D-02, "no colon no dash") before being adopted — both have identical shape under any content-only rule. Nesting depth (via indentWidth) is the actual distinguishing signal: a bullet indented deeper than the currently-open decision's own bullet is elaboration, folded into its text like a continuation line, never tested against the parse-miss guard. A bullet at the same-or-shallower indent is unchanged. Scope-narrowing disclosed, not silent: of the ADR's four named bugs (#3197, #3169, #2570, #2528), three no longer need this phase's work. were independently fixed and closed since the ADR was authored — #2570's fix is already a correctly-bounded regex per the ADR's own decidability test (no scanner needed); #2528's fix is a deliberate, twice-reviewed non-scanner design (its own code comment records a scanner-based attempt that regressed a symmetric case and was reverted) that this phase does not disturb. Only #3169 required new work. get_impact: isGitSubcommand CRITICAL/196 affected symbols, parseDecisionLines CRITICAL/164 affected symbols (ADR §6 due diligence). Six-gate ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary. Design: .gsd/phase/chore-3414-tokenizer-first-seam/40-design.md Test matrix: .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3414): add required fast-check property tests per code review TESTING-STANDARDS.md:169 requires at least one fast-check property test for any module that implements parsing — src/token-scanner.cts had none, an orthogonal Standards-axis review finding. Adds two seeded property tests (mirroring Phase 1/2's fast-check-setup.cjs convention): indentWidth counts exactly a generated leading-space run; tokenizeShellLike round-trips a generated array of whitespace/quote-free words joined with single spaces. The design doc's own "no property test needed" rationale was wrong — it argued no algebraic law applied, but the standard is unconditional for parsing modules regardless of whether one "feels" applicable. Corrected in .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md. Also fixes two Spec-axis wording drifts the same review found between the design doc and the shipped code (doc-only, no behavior change): extractBranchArgument's documented signature dropped an unused subVariants parameter that was never implemented, and the #3169 fail-first fixture description corrected from "15-decision plan via cmdDecisionCoverageVerify" to the actual compact 3-decision analog via the real blocking gate, check.decision-coverage-plan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3414): add changeset for #3169 fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3414): backfill changeset pr number to 3424 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
185da024cb |
fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block (#2881)
* test(#2770): empty contextPath argument must fail closed, not green-skip the decision-coverage gate The handler conflated empty-arg (caller error) with file-missing (legitimate skip), returning passed:true/skipped on an empty argument. Add: empty arg → passed:false; real-path-to-absent-file → legitimate green skip preserved; omitted arg → fail closed. * fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block Handler (check-command-router.cts): split the guard — empty/missing contextPath argument is a caller error (fail closed, passed:false, mirrors #1365); a real path whose file genuinely does not exist keeps the legitimate green skip. Workflow (plan-phase.md): recompute CONTEXT_PATH inside the consuming Bash block (it was set in the step-1 init block, which does not survive into the separately- spawned gate block — so the gate ran with an empty arg and silently green-skipped). * chore(#2770): changeset fragment * fix(#2770): guard workflow empty-glob case (review blocker) + update drift-guard test The handler now fails closed on an empty contextPath arg, so the workflow's unguarded glob (empty when a phase genuinely has no CONTEXT.md) would invoke the gate with an empty arg → passed:false → exit 1, hard-halting the legitimate 'Continue without context' plan-phase path. Guard the empty-glob case: only run the gate when a CONTEXT.md actually exists. Update the F1 drift-guard test (which gave false coverage — it only checked for the ${CONTEXT_PATH} token) to assert the in-block recompute AND the empty-glob guard. * fix(#2770): keep plan-phase.md under ADR-857 size cap + ack emitted drift + fix drift-guard window The workflow fix grew plan-phase.md past the ADR-857 phase-6 size cap (94519B) and triggered emitted-attribution. Condense adjacent §13a prose/JSON to offset (net +89B, under cap). Add tests/emitted-drift-ack.json acknowledging the residual growth. Widen the drift-guard test window (the gate invocation is now nested in the empty-glob guard, so the old 400-char window missed the glob recompute). * chore(#2770): backfill changeset PR number (2881) --------- Co-authored-by: Test <test@example.com> |
||
|
|
517bae8d6d |
fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message Bug: check.decision-coverage-plan's remediation message told the user to cite decisions "(or body)" but extractPlanDesignatedSections only scanned <objective>/<tasks>/<task>/<action>. A decision cited in <read_first>, <behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to the gate — false BLOCKING coverage gap, plus the message's own fix-hint sent the user to "the body" where re-citing still failed. Two-part fix (must change together — that drift was the bug): 1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also match <read_first>, <behavior>, <verify>, <acceptance_criteria>, <done>. These are all planner-canonical tags the planner is told to use (plan-phase.md:830-862, plan-phase.md:772). The body negative- lookahead mirrors the opening-tag set so each tag's body is captured independently. 2. Correct buildPlanMessage to name ONLY the surfaces the extractor actually scans (front-matter must_haves/truths/objective, designated markdown headings, and the nine planner-canonical tag bodies). The misleading "(or body)" clause is gone. Also updates the planner's documented contract (agents/gsd-planner.md:69) and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to reflect the wider scan. Regression tests in tests/decisions.test.cjs cover each newly-scanned tag body, a control (no citation still uncovered), and a message/extractor parity assertion that names every scanned surface — so the two cannot drift apart again. Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage is a separate command (decision-coverage-verify) checking shipped artifacts, not plan citations — untouched. * chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision- coverage contract (5 new scanned tag names + heading clarification). Growth is justified: the contract surface is itself the fix — the prior text under-described what the gate scans, which was the bug. Updates: - tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294) - 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime) - tests/fixtures/install-tree/*.json (regenerated by gen:golden) * fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting Code review (subagent) flagged a Medium edge-case regression from the single-alternation regex: when a newly-scanned tag nests inside another scanned tag, the alternation's negative lookahead halts the outer tag's body at the inner tag — losing any D-NN citation in the outer tag's prefix prose. Concretely: <action>per D-05 <verify>npm test</verify></action> → 3-tag alternation (old): captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught → 9-tag alternation (bug): captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST Switches extractXmlTagBodies to per-tag matching: each tag gets its own regex whose negative-lookahead tempers only against the SAME tag's reopening. So <verify> inside <action> is absorbed into <action>'s body (D-05 caught) AND <verify> is matched separately on its own pass. Per-tag preserves both: - the reporter's case (sibling tags inside <read_first>) - nested-tag citations in outer-tag prefix prose - ReDoS safety (each per-tag regex keeps the #2128 body tempering) Also adds the reviewer's other requested edge-case tests: - non-scanned tag (<name>) bearing D-NN must NOT count - self-closing form <read_first /> safely ignored - attribute form <verify type="...">D-NN</verify> (canonical planner shape) - CRLF newlines in tag body do not break capture * chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md |
||
|
|
f15eb5f5c9 |
fix(#2347): make the decision-shape evidence test format-agnostic (#2389)
#1365's fail-loud guard reused the parser's own D- grammar as its evidence test, so a populated <decisions> block using any other ID prefix (e.g. D5-01) was invisible to both parser and guard, collapsing could-not-parse into a clean none-present pass. Add an ID-shaped bold-lead-in probe as format-agnostic evidence on both parse paths; empty/prose scaffolds stay none-present. Graduates the #2371 d5-prefix representative fixture to its expected* assertion. Closes #2347. Admin-merged (self-review bypass) with full green CI. |
||
|
|
b205e4c2b2 |
fix(#1639): parseDecisions handles the titled-colon bullet form (#1665)
* fix(#1639): parseDecisions handles titled-colon bullet form bulletColonRe anchors on ':**' (colon immediately before close-bold) and bulletEmDashRe requires an em-dash, so the titled-colon form '- **D-NN: Title.** body' (title between the colon and the closing **) matched neither and was dropped by the parse-miss guard. When all decisions used the titled convention, parseDecisions returned 0 and check.decision-coverage- plan passed vacuously — the same false-coverage failure mode as #1343/#1364/#1365. Add a third per-form regex bulletTitledColonRe, checked LAST (strict superset of bulletColonRe, so it only catches bullets the other two miss — minimal blast radius); id + [tags] trackability honored. Regression folded into decisions.test.cjs: titled-colon parses, coexists with colon/em-dash, tags, all-titled-13 no longer vacuously 0. * fix(#1639): tighten titled-colon title to [^:*]* so malformed pre-colon-run bullets still reject The first cut's title run [^*]* was too permissive: it matched a genuinely-malformed bullet with a colon in the pre-separator freeform run (e.g. 'D-07 ratio 3:1:**') by treating the 3:1 colon as the separator, regressing the #1343 parse-miss guard tests. Tighten the title to [^:*]* (no colon, no star) so the separator colon remains the only colon permitted before ** — matching bulletColonRe's existing [^:*]* discipline. Valid titled forms (colon-free titles) still parse; the malformed colon-in-freeform case still falls through to the parse-miss guard. * chore(#1639): backfill changeset pr ref to 1665 |
||
|
|
e58b5e1721 |
fix(#1364): decisions adopt markdown-sectionizer seam + fail-loud coverage gate (epic #1372 T1) (#1386)
* test(#1364,#1365): add decisions regression tests (fail-first proof) Adds tests/decisions.test.cjs with: - #1364 recall tests: parseDecisions from markdown-header + em-dash bullets (these FAIL on pre-T1 code, proving the bug is present before the fix) - #1365 fail-loud tests: check.decision-coverage-plan must return passed:false for decision-shaped but 0-extracted content (FAIL pre-T1, gate silently passed) - extractDecisions outcome enum tests (could-not-parse/none-present/parsed) - Parser QA matrix: CRLF, unicode headings, fenced-code suppression, both bullet forms - Boundary/threshold tests at limit-1 (0), limit (1) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1364,#1365): adopt markdown-sectionizer seam in decisions.cts; add fail-loud gate #1364 — Recall: decisions.cts now uses the seam's extractTaggedBlocks and collectSection for the markdown-header fallback path. Em-dash bullet form (- **D-NN — title** body) is now recognised alongside the existing colon form. #1365 — Fail-loud: adds extractDecisions() returning a typed DecisionExtraction { decisions, outcome } where outcome is 'parsed' | 'none-present' | 'could-not-parse'. The blocking gate (cmdDecisionCoveragePlan) now treats could-not-parse as passed:false with a format-mismatch reason instead of the prior silent passed:true/skip. gap-checker runGapAnalysis surfaces 'extracted 0 of N — possible format mismatch' for could-not-parse instead of 'No requirements or decisions to check'. parseDecisions remains a thin delegate over extractDecisions, so all existing callers are unaffected. Seam adoption: stripFencedCode (seam), extractTaggedBlocks(content,'decisions') (seam), collectSection(content, /decisions?/i, {levelBounded,stripFences}) (seam). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1364,#1365): tighten could-not-parse, parse-miss fail-loud, curly-quote discretion, gap-checker FIX D FIX A: empty <decisions> scaffolds and all-prose sections no longer return could-not-parse; outcome is none-present unless the block/section contains a \bD- token or a parse-miss, preventing false blocks on legitimate phases. FIX B: parseDecisionLines now tracks parse-misses (D-NN-shaped bullets that fail both regexes); extractDecisions returns could-not-parse when parseMisses>0 even if some decisions parsed — silent drops no longer mask format errors. FIX C: curly-quote normalization regex now includes actual U+2018/U+2019 characters so '### Claude's Discretion' (curly apostrophe) correctly yields trackable:false (regression vs pre-T1 behavior). FIX D: gap-checker runGapAnalysis surfaces the decision could-not-parse format-mismatch signal independently of whether requirements items exist — previously masked inside `if (items.length === 0)`. Adds 14 behavioral regression tests (fail-first verified manually before fixes). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1365): fail-loud gate on parse-miss regardless of covered decisions Change the `could-not-parse` guard in `cmdDecisionCoveragePlan` and `cmdDecisionCoverageVerify` from `decisions.length === 0 && outcome === 'could-not-parse'` to fire on `outcome === 'could-not-parse'` alone. Previously a CONTEXT.md with a valid D-01 (covered by the plan) plus a malformed D-02 (parse-miss) would skip the guard (length === 1), proceed to coverage, find D-01 covered, and silently return passed:true — hiding the D-02 parse-miss entirely. Adds a gate-level fail-first test that places D-01 into a ## Must Haves section (DESIGNATED_HEADINGS_RE match) so coverage of D-01 would pass on its own, proving the only path to passed:false is the parse-miss fix. Also adds the matching verify-side advisory assertion. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(#1364,#1365): add Fixed changeset (pr:0 placeholder) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1364): backfill changeset PR number (1386) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |