Files
msd-core/tests/nsegment-phase-grammar.test.cjs
0xdhx 092d9256b8 fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it (#4768)
* test(#4748): pin the letter-axis defect at the seven shell sites outside #4660's six

Extends tests/nsegment-phase-grammar.test.cjs one class over: for each of the
seven sites the live shell lines are read off disk by anchor and executed in
bash against a letter-suffixed fixture. The four `$((10#$PHASE_INT))` split
sites must yield PHASE_N without a shell error for `03A` / `12A` / `3A` /
`03A.1.2` and the commit-scope ERE they build must match both `feat(3A-01):`
and `feat(03A-1):`; the review-file lookup must bind init's `padded_phase`
rather than re-pad in shell; the `--from`/`--to`/`--only` and
plan-review-convergence extractions must return `12A` / `23A.1.2` (and
`23.1.2`) whole; the legacy normalizer must pad `3A` to `03A` and must not
mangle an already-padded `08`. Every pre-existing shape (`06`, `08.5`,
`23.1.2`, `36.14`) is a regression control.

tests/init.test.cjs asserts `init execute-phase` emits `padded_phase` for a
directory-backed `03A`, a ROADMAP-only `4B` (→ `04B`), the existing ROADMAP
fallback `1` (→ `01`), and `null` when the phase is not found.

Negative control against the unfixed tree: 41 failures in the grammar file,
exactly the "(fails before the fix)" cases and the three derived from them
(scope ERE, three-flag extraction, the `08` octal trap); 2 in init.test.cjs,
both the new assertions. Every regression control already green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it

The canonical phase-number grammar (src/phase-id.cts) is digits, an optional
uppercase letter, then dotted segments — `12A`, `3A`, `23A.1.2` are documented
shapes that `init`, `phase-id.cts` and `phase remove` renumbering already
round-trip. Seven shell sites in shipped workflows and references still
assumed digits-and-dots. Four classes, one fix each:

Class 1 — `PHASE_INT=${PHASE_NUMBER%%.*}; $((10#$PHASE_INT))` (execute-phase.md
×2, completion-reconciliation.md, tdd.md). The post-#4619 split stops at the
first DOT, so on `03A` the "integer" is `03A` and bash aborts with `value too
great for base`. Split at the first NON-DIGIT instead (`%%[!0-9]*`): the
integer half is a pure digit run, and the letter rides along in the rest the
way the dotted fraction already did — `03A.1.2` → PHASE_N `3A\.1\.2`, so the
#4003 zero-pad-tolerant scope ERE matches both `feat(3A-01):` and
`feat(03A-1):`. Byte-identical output for every id that worked before.

Class 2 — `PADDED=$(printf "%02d" "${PHASE_NUMBER}")` before the REVIEW.md
lookup (execute-phase.md). `printf` cannot pad a letter id (prints `03`,
exits 1) — and cannot even re-pad an already-padded `08`, which bash reads as
an invalid octal and prints as `00`, so the lookup resolved phases 08 and 09
to `00-REVIEW.md` today. The disk path hands the workflow the directory's
padded number but the ROADMAP fallback hands it the heading's bare one, which
is why the re-pad existed. `cmdInitExecutePhase` now emits `padded_phase`
through `normalizePhaseName`, exactly as the plan-phase and code-review inits
do, and the workflow binds `{padded_phase}` instead of re-deriving.

Class 3 — `grep -oE '[0-9]+\.?[0-9]*'` (autonomous.md `--from`/`--to`/`--only`,
plan-review-convergence.md). Stops at the letter, so `--from 12A` ran from
phase 12 with no error. Now the canonical ERE `[0-9]+[A-Z]?(\.[0-9]+)*`, which
also closes the single-segment dot-axis gap the same shape carried (`23.1.2`
→ `23.1`, #4568's class in a spelling neither lint saw).

Class 4 — the legacy manual normalizer (phase-argument-parsing.md, reached
from mvp-phase.md). Its two branches (`^[0-9]+$`, `^[0-9]+\.[0-9]+$`) left
`12A` unpadded and never padded `3A` to the `03A` a directory carries; its
integer branch also hit the same `printf` octal trap on `08`. One branch for
the whole canonical token now, padding the digit run via `$((10#…))`.
Whether this legacy surface should instead be retired in favour of `init`'s
normalization is the maintainer call the issue names; extending it keeps the
documented contract true either way.

Driven end to end: `init execute-phase 3A` on a fixture with a
`03A-letter-variant/` directory emits `phase_number: "03A"` and now
`padded_phase: "03A"`; on a ROADMAP-only `### Phase 4B:` it emits `"4B"` /
`"04B"`. The issue's own evidence line claimed `padded_phase` was already in
the execute-phase init output — it was not; that key is emitted by the
code-review / plan-phase inits, which is where the claim was read from.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): extend lint-phase-id-drift with three ratchets for letter-hostile phase-id consumers

The rules that landed with #4619, #4568 and #4660 police grammar MIRRORS —
regexes that describe a phase id. The #4748 sites are CONSUMERS of one, and
every existing rule reported clean on them: the shell-arithmetic rule's
`_INT` escape trusts a NAME the dot-only split did not earn on `03A`; the
`[0-9]+\.?[0-9]*` shape is neither the bounded form the single-segment rule
bans nor the unbounded form the letterless rule inspects; and nothing looked
at `printf "%02d"` at all. Three narrow additions, one per shape:

- findDotOnlyIntegerSplitDrift — `X_INT=${<phase-var>%%.*}`; the safe split
  is `%%[!0-9]*`. Keys on the SOURCE variable being phase-carrying.
- findLooseDottedPhaseRegexDrift — `[0-9]+\.?[0-9]*` / `\d+\.?\d*` on a
  phase-carrying line; the canonical form is `[0-9]+[A-Z]?(\.[0-9]+)*`.
  Disjoint from the two sibling regex rules by construction.
- findShellPhasePrintfPadDrift — `printf "%0Nd" …` whose arguments name a
  phase-carrying, non-`_INT` variable; a pad of an `_INT` via `$((10#…))`
  and a `{padded_phase}` binding are the sanctioned shapes.

Same `<!-- phase-id-owner: … -->` sanction, same scan roots as their nearest
sibling (shell idioms over workflows + references, the regex shape over
workflows + references + agents), same documented limit of a per-line
textual scan. The post-#4619 comment that described the `_INT` convention
as proven by `%%.*` is corrected to name the digit-run split. Confirmed
against the base commit: each rule fires on exactly its own unfixed sites
(2+1+1, 3+1, 1+1) and zero violations remain on the fixed tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* docs(#4748): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline and acknowledge emitted growth

The three top-level workflow files below grew by the letter-aware split, the
canonical extraction ERE, the `{padded_phase}` binding, and the comment lines
that name the grammar each site now honours. The committed compact-content
benchmark moved with them; refreshed with `benchmark-compact-content.cjs
--write` (aggregate reduction 15.47% -> 15.45%).

Emitted-Drift-Ack-Growth: execute-phase.md — #4748: first-non-digit PHASE_INT split at the plan-selection and TDD-gate sites, `{padded_phase}` binding at the REVIEW.md lookup, and the comments naming why (482 bytes)
Emitted-Drift-Ack-Growth: autonomous.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the --from/--to/--only extractions plus one comment naming the grammar (249 bytes)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the phase extraction plus one comment naming the grammar (160 bytes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): name padded_phase in execute-phase.md's init parse list

A `{field}` token inside a workflow bash block is substituted from the init
JSON only for fields the workflow tells the model to parse. `phase_number`
is on that list; `padded_phase` was not, so the `PADDED="{padded_phase}"`
binding at the review lookup would have been a literal — for every phase,
not only letter ones. Found by the pre-file adversarial review (claim 2, the
author's own named suspicion); the test now asserts the parse list carries
the field beside `phase_number`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on its source and widen the printf rule to any %d form

Two false negatives from the pre-file adversarial review of the three #4748
ratchets: `PHASE_PREFIX=${PHASE_NUMBER%%.*}` escaped the split rule because
the destination did not end in `_INT` (the defect is the split, not the
name it lands in), and `printf '%02d'` / `printf "%2d"` escaped the printf
rule because it required double quotes and the zero flag (`%d` cannot parse
a letter id under any width). Both rules now key on the phase-carrying
SOURCE alone; base-site firing counts are unchanged (2+1+1, 1+1) and the
fixed tree stays at zero. The `[[:digit:]]` spelling and the `/phase/i`
heuristic remain the sibling rules' documented limits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline after the parse-list edit

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): compose init's emitted padded_phase through the live REVIEW.md lookup

The Class 2 site is a `{padded_phase}` template token, which no test can
execute as written. This substitutes the value init emits
(`normalizePhaseName`) into the three live lookup lines and runs them
against a fixture, so the emitted value, the binding, the path construction
and the status extraction are exercised together — `03A-REVIEW.md` and
`08-REVIEW.md` each resolve to their own status. Suggested by the resumed
adversarial review pass (claim C).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): move the #4619 and #4003 source-parity pins to the letter-safe split

tests/execute-phase-decimal-arithmetic.test.cjs and
tests/safe-resume-gate-anchoring.test.cjs pin the four Class 1 sites'
snippet byte-for-byte, so the first-non-digit split reddened both in the
whole-suite run (scripts/ci-test-scope.cjs does not select either file for
a workflow edit — the scoped run was green). The pinned snippet is now the
shipped one, and the behavioural half of the #4619 file gains the letter
case (`03A` → `3A`, `23A.1.2` → `23A\.1\.2`) beside its decimal cases.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on the _INT destination again, tolerating the quoted spelling

Keying on the source alone (the previous commit's widening, from a review
probe) flags `PARENT_PHASE="${PHASE_NUMBER%%.*}"` in
gap-closure-artifacts.md — a correct derivation that wants everything
before the first dot, letter included. The defect this rule polices is a
dot split INTO the name the shell-arithmetic rule trusts as an integer, so
`_INT` is the discriminator on purpose; the quoted spelling that site uses
is now tolerated so the same shape into an `_INT` cannot hide behind it.
Base-site firing unchanged (2+1+1), zero on the fixed tree, and the
parent-phase line is pinned as a silent case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): use t.after() for the composition test's fixture cleanup

CONTRIBUTING forbids try/finally inside a test body; the per-test cleanup
form is `t.after(() => cleanup(dir))`. Flagged by the filing driver's
test-ruleset gate before the PR was created.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): set changeset fragment pr to 4768

* chore(#4748): refresh the compact-content benchmark baseline after rebasing onto next

Regenerated with `node scripts/benchmark-compact-content.cjs --write` on the
rebased tree (base 0d6bf19bf); `--check` confirms it matches the live recompute.
Only the execute-phase split and the aggregate totals differ from next's copy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 23:31:36 -04:00

609 lines
27 KiB
JavaScript

'use strict';
/**
* #4568 (epic #4634) — six shell snippets embedded in workflow/agent markdown
* validate or extract phase numbers with the regex shape `[0-9]+(\.[0-9]+)?`
* (or its `\d` near-variant) — an optional SINGLE dotted segment. Any
* three-or-more-segment phase id (e.g. `23.1.2`, produced by a nested `phase
* insert`) is either hard-rejected or silently truncated to the wrong value.
* The canonical grammar in src/phase-id.cts already uses the unbounded form
* (`\d+(?:\.\d+)*`) — shell cannot import that module, so the fix is textual
* parity: widen `?` to `*` at each site.
*
* These tests are BEHAVIORAL: for each site, the actual regex/extraction
* line is read live off disk (via a narrow, anchored string search) and
* executed in a real bash subprocess — never hand-retyped — so the test
* breaks loudly if a future edit changes a site's shape instead of silently
* drifting from the real file.
*/
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { execFileSync } = require('node:child_process');
const TIMEOUT = 5000;
const CODE_REVIEW = path.join(__dirname, '..', 'gsd-core', 'workflows', 'code-review.md');
const CODE_REVIEW_FIX = path.join(__dirname, '..', 'gsd-core', 'workflows', 'code-review-fix.md');
const CODE_FIXER = path.join(__dirname, '..', 'agents', 'gsd-code-fixer.md');
const CODE_FIXER_COMPACT = path.join(__dirname, '..', 'agents', 'gsd-code-fixer.compact.md');
const EXECUTE_PLAN = path.join(__dirname, '..', 'gsd-core', 'workflows', 'execute-plan.md');
const PLAN_PHASE = path.join(__dirname, '..', 'gsd-core', 'workflows', 'plan-phase.md');
/**
* Pure: find the line containing `anchor` and pull the regex substring
* between `=~ ` and ` ]]` on it. Throws loudly if either the anchor or the
* pattern shape is not found, so a future rewrite of the site's surrounding
* code breaks this test instead of silently testing stale text.
*/
function extractAnchoredRegex(fileText, anchor) {
const lines = fileText.split('\n');
const line = lines.find((l) => l.includes(anchor));
assert.ok(line, `anchor not found: ${anchor}`);
const m = line.match(/=~\s+(\S+)\s+\]\]/);
assert.ok(m, `no "=~ <pattern> ]]" shape found on anchor line: ${line}`);
return m[1];
}
/**
* Pure: find the line containing `anchor` and pull the single-quoted
* `grep -oE '...'` pattern off it.
*/
function extractGrepPattern(fileText, anchor) {
const lines = fileText.split('\n');
const line = lines.find((l) => l.includes(anchor));
assert.ok(line, `anchor not found: ${anchor}`);
const m = line.match(/grep -oE '([^']+)'/);
assert.ok(m, `no grep -oE '...' shape found on anchor line: ${line}`);
return m[1];
}
/** Run a validating-site regex (bash `[[ =~ ]]`) against `value`, returning true/false. */
function matchesValidatingRegex(pattern, value) {
const script = `if [[ "$TEST_INPUT" =~ ${pattern} ]]; then echo MATCH; else echo NOMATCH; fi`;
const out = execFileSync('bash', [], {
input: script,
encoding: 'utf8',
timeout: TIMEOUT,
env: { ...process.env, TEST_INPUT: value },
}).trim();
return out === 'MATCH';
}
describe('#4568 — validating sites accept N-segment phase ids and still reject injection', () => {
const sites = [
{ name: 'code-review.md', file: CODE_REVIEW, anchor: 'if ! [[ "$PADDED_PHASE" =~ ' },
{ name: 'code-review-fix.md', file: CODE_REVIEW_FIX, anchor: 'if ! [[ "$PADDED_PHASE" =~ ' },
{ name: 'gsd-code-fixer.md', file: CODE_FIXER, anchor: 'if ! [[ "$padded_phase" =~ ' },
{ name: 'gsd-code-fixer.compact.md', file: CODE_FIXER_COMPACT, anchor: 'if ! [[ "$padded_phase" =~ ' },
];
for (const site of sites) {
describe(site.name, () => {
const text = fs.readFileSync(site.file, 'utf8');
const pattern = extractAnchoredRegex(text, site.anchor);
test('regression control: 1-segment id (6) matches', () => {
assert.equal(matchesValidatingRegex(pattern, '6'), true);
});
test('regression control: 2-segment id (36.14) matches', () => {
assert.equal(matchesValidatingRegex(pattern, '36.14'), true);
});
test('N-segment id (23.1.2) matches (fails before the fix)', () => {
assert.equal(matchesValidatingRegex(pattern, '23.1.2'), true);
});
test('path-traversal injection (../1) is rejected', () => {
assert.equal(matchesValidatingRegex(pattern, '../1'), false);
});
test('shell-metacharacter injection (1; rm -rf /) is rejected', () => {
assert.equal(matchesValidatingRegex(pattern, '1; rm -rf /'), false);
});
test('empty string is rejected', () => {
assert.equal(matchesValidatingRegex(pattern, ''), false);
});
});
}
});
describe('#4568 — execute-plan.md extracts the full N-segment phase from a plan filename', () => {
const text = fs.readFileSync(EXECUTE_PLAN, 'utf8');
const pattern = extractGrepPattern(text, 'grep -oE');
function extractPhase(planPath) {
const script = `echo "$PLAN_PATH" | grep -oE '${pattern}'`;
let out;
try {
out = execFileSync('bash', [], {
input: script,
encoding: 'utf8',
timeout: TIMEOUT,
env: { ...process.env, PLAN_PATH: planPath },
}).trim();
} catch {
out = '';
}
return out;
}
test('regression control: 1-segment plan filename extracts correctly', () => {
assert.equal(extractPhase('/x/06-01-PLAN.md'), '06-01');
});
test('regression control: 2-segment plan filename extracts correctly', () => {
assert.equal(extractPhase('/x/36.14-01-PLAN.md'), '36.14-01');
});
test('N-segment plan filename extracts the FULL phase, not a truncated one (fails before the fix)', () => {
assert.equal(extractPhase('/x/23.1.2-01-PLAN.md'), '23.1.2-01');
});
});
describe('#4568 — plan-phase.md captures the full N-segment --research-phase value', () => {
const text = fs.readFileSync(PLAN_PHASE, 'utf8');
const pattern = extractAnchoredRegex(text, '=~ --research-phase[[:space:]]+(');
function captureResearchPhase(args) {
const script = [
'if [[ "$ARGUMENTS" =~ ' + pattern + ' ]]; then',
' echo "${BASH_REMATCH[1]}"',
'else',
' echo NOMATCH',
'fi',
].join('\n');
return execFileSync('bash', [], {
input: script,
encoding: 'utf8',
timeout: TIMEOUT,
env: { ...process.env, ARGUMENTS: args },
}).trim();
}
test('regression control: 1-segment --research-phase captures correctly', () => {
assert.equal(captureResearchPhase('--research-phase 6'), '6');
});
test('regression control: 2-segment --research-phase captures correctly', () => {
assert.equal(captureResearchPhase('--research-phase 36.14'), '36.14');
});
test('N-segment --research-phase captures the FULL value, not a truncated one (fails before the fix)', () => {
assert.equal(captureResearchPhase('--research-phase 23.1.2'), '23.1.2');
});
});
// ---------------------------------------------------------------------------
// #4660 — the LETTER axis. #4568 widened the six sites on the segment-count
// axis only; the canonical grammar also admits an optional single uppercase
// letter after the leading digits (`12A`, `3A`, `23A.1.2` — a documented
// phase-number shape in CONFIGURATION.md, relied on by renameIntegerPhases in
// src/phase.cts). These tests prove each site's live pattern and the canonical
// source AGREE on that axis, in both directions, rather than each merely
// "looking right" in isolation.
// ---------------------------------------------------------------------------
// The canonical grammar is read from the committed bin/lib mirror the other
// grammar tests use (shell cannot import it; the test can).
const { PHASE_NUMBER_TOKEN_SOURCE } = require('../gsd-core/bin/lib/phase-id.cjs');
const { splitLines } = require('../gsd-core/bin/lib/text-lines.cjs');
const CANONICAL_ANCHORED = new RegExp('^(?:' + PHASE_NUMBER_TOKEN_SOURCE + ')$');
// Inputs the canonical grammar ACCEPTS. `03A` is what `normalizePhaseName('3A')`
// emits, i.e. the real `padded_phase` the four validating sites receive from
// `init`; the bare forms are what a user types or names a directory with.
const LETTER_ACCEPT = ['12A', '3A', '03A', '23A.1.2'];
// Inputs the canonical grammar REJECTS on the same axis — a parity test that
// only checks accepts would pass against `.*`. Lowercase is refused because
// the canonical source is case-sensitive `[A-Z]` (the case-flexible variant
// is a separate, deliberately distinct axis — see phase-id.cts).
const LETTER_REJECT = ['3a', '3AB', 'A3', '3A.', '3.A', '3A-1'];
describe('#4660 — canonical grammar fixture agrees with the inputs this file uses', () => {
for (const v of LETTER_ACCEPT) {
test(`canonical accepts ${v}`, () => {
assert.equal(CANONICAL_ANCHORED.test(v), true);
});
}
for (const v of LETTER_REJECT) {
test(`canonical rejects ${v}`, () => {
assert.equal(CANONICAL_ANCHORED.test(v), false);
});
}
});
describe('#4660 — validating sites agree with the canonical grammar on the letter axis', () => {
const sites = [
{ name: 'code-review.md', file: CODE_REVIEW, anchor: 'if ! [[ "$PADDED_PHASE" =~ ' },
{ name: 'code-review-fix.md', file: CODE_REVIEW_FIX, anchor: 'if ! [[ "$PADDED_PHASE" =~ ' },
{ name: 'gsd-code-fixer.md', file: CODE_FIXER, anchor: 'if ! [[ "$padded_phase" =~ ' },
{ name: 'gsd-code-fixer.compact.md', file: CODE_FIXER_COMPACT, anchor: 'if ! [[ "$padded_phase" =~ ' },
];
for (const site of sites) {
describe(site.name, () => {
const text = fs.readFileSync(site.file, 'utf8');
const pattern = extractAnchoredRegex(text, site.anchor);
for (const v of LETTER_ACCEPT) {
test(`letter-suffixed id ${v} matches (fails before the fix)`, () => {
assert.equal(matchesValidatingRegex(pattern, v), true);
assert.equal(matchesValidatingRegex(pattern, v), CANONICAL_ANCHORED.test(v));
});
}
for (const v of LETTER_REJECT) {
test(`canonical-invalid ${v} is still rejected (parity, not a blanket widening)`, () => {
assert.equal(matchesValidatingRegex(pattern, v), false);
assert.equal(matchesValidatingRegex(pattern, v), CANONICAL_ANCHORED.test(v));
});
}
});
}
});
describe('#4660 — execute-plan.md extracts the full letter-suffixed phase from a plan filename', () => {
const text = fs.readFileSync(EXECUTE_PLAN, 'utf8');
const pattern = extractGrepPattern(text, 'grep -oE');
function extractPhase(planPath) {
const script = `echo "$PLAN_PATH" | grep -oE '${pattern}'`;
let out;
try {
out = execFileSync('bash', [], {
input: script,
encoding: 'utf8',
timeout: TIMEOUT,
env: { ...process.env, PLAN_PATH: planPath },
}).trim();
} catch {
out = '';
}
return out;
}
// Before the fix a letter-suffixed filename either extracts NOTHING (the
// digit run is followed by the letter, so `-[0-9]+` never attaches) or the
// wrong tail (`23A.1.2-01` → `1.2-01`). Both are silent mis-extractions.
for (const [planPath, expected] of [
['/x/12A-01-PLAN.md', '12A-01'],
['/x/03A-02-PLAN.md', '03A-02'],
['/x/23A.1.2-01-PLAN.md', '23A.1.2-01'],
]) {
test(`${planPath} extracts ${expected} (fails before the fix)`, () => {
const got = extractPhase(planPath);
assert.equal(got, expected);
// The phase half of the extraction is canonical-valid — parity with src/phase-id.cts.
assert.equal(CANONICAL_ANCHORED.test(got.replace(/-\d+$/, '')), true);
});
}
test('regression control: the letter class is admitted at the PHASE position only (plan numbers stay digit-only)', () => {
// `12A-B1`: the plan half must start with a digit, so nothing attaches to
// `12A-` and the digit-only tail `1` has no `-[0-9]+` after it either.
assert.equal(extractPhase('/x/12A-B1-PLAN.md'), '');
});
});
describe('#4660 — plan-phase.md captures the full letter-suffixed --research-phase value', () => {
const text = fs.readFileSync(PLAN_PHASE, 'utf8');
const pattern = extractAnchoredRegex(text, '=~ --research-phase[[:space:]]+(');
function captureResearchPhase(args) {
const script = [
'if [[ "$ARGUMENTS" =~ ' + pattern + ' ]]; then',
' echo "${BASH_REMATCH[1]}"',
'else',
' echo NOMATCH',
'fi',
].join('\n');
return execFileSync('bash', [], {
input: script,
encoding: 'utf8',
timeout: TIMEOUT,
env: { ...process.env, ARGUMENTS: args },
}).trim();
}
// Before the fix the capture stops at the digit boundary: `12A` → `12`.
for (const v of ['12A', '3A', '23A.1.2']) {
test(`--research-phase ${v} captures ${v}, not its digit prefix (fails before the fix)`, () => {
const got = captureResearchPhase(`--research-phase ${v}`);
assert.equal(got, v);
assert.equal(CANONICAL_ANCHORED.test(got), true);
});
}
});
// ---------------------------------------------------------------------------
// #4748 — the letter axis at the seven shell sites OUTSIDE #4660's six. These
// are not grammar mirrors but consumers of the id: the post-#4619
// `PHASE_INT=${PHASE_NUMBER%%.*}; $((10#$PHASE_INT))` split (four sites),
// the `printf "%02d"` re-pad before the REVIEW.md lookup (one site), the
// `[0-9]+\.?[0-9]*` argument extraction (two files, four lines) and the
// legacy manual normalizer. On a letter-suffixed id the first aborts bash,
// the second prints the wrong file, the last two silently truncate. Same
// discipline as above: each site's live lines are read off disk by anchor
// and executed in a real bash subprocess.
// ---------------------------------------------------------------------------
const EXECUTE_PHASE = path.join(__dirname, '..', 'gsd-core', 'workflows', 'execute-phase.md');
const COMPLETION_RECONCILIATION = path.join(
__dirname, '..', 'gsd-core', 'workflows', 'execute-phase', 'steps', 'completion-reconciliation.md',
);
const TDD_REF = path.join(__dirname, '..', 'gsd-core', 'references', 'tdd.md');
const AUTONOMOUS = path.join(__dirname, '..', 'gsd-core', 'workflows', 'autonomous.md');
const PLAN_REVIEW_CONVERGENCE = path.join(__dirname, '..', 'gsd-core', 'workflows', 'plan-review-convergence.md');
const PHASE_ARGUMENT_PARSING = path.join(__dirname, '..', 'gsd-core', 'references', 'phase-argument-parsing.md');
/**
* Pure: the indexes of every line containing `anchor`. Asserts the count so a
* site that is added, removed or renamed breaks this test loudly instead of
* silently narrowing what it covers (execute-phase.md carries the split TWICE
* — plan selection and the TDD gate — and both must stay under test).
*/
function findAnchoredLineIndexes(lines, anchor, expectedCount) {
const idx = [];
lines.forEach((l, i) => {
if (l.includes(anchor)) idx.push(i);
});
assert.equal(
idx.length,
expectedCount,
`expected ${expectedCount} line(s) containing ${JSON.stringify(anchor)}, found ${idx.length}`,
);
return idx;
}
/** Run `script` in bash with `env` merged in; never throws — returns { status, stdout, stderr }. */
function runBash(script, env) {
try {
const stdout = execFileSync('bash', [], {
input: script,
encoding: 'utf8',
timeout: TIMEOUT,
env: { ...process.env, ...env },
stdio: ['pipe', 'pipe', 'pipe'],
});
return { status: 0, stdout: stdout.trim(), stderr: '' };
} catch (e) {
return { status: e.status, stdout: String(e.stdout || '').trim(), stderr: String(e.stderr || '').trim() };
}
}
// What each Class 1 site must compute from the id it is handed: the integer
// half zero-stripped for the anchored `0*` commit-scope ERE, everything after
// it carried through with dots escaped. `03A` is the padded form `init` emits
// for a `03A-slug/` directory; `12A` / `3A` are the bare forms; `03A.1.2` is
// the letter-and-N-segment combination the canonical grammar admits.
const CLASS1_CASES = [
// [PHASE_NUMBER, expected PHASE_N]
['03A', '3A'],
['12A', '12A'],
['3A', '3A'],
['03A.1.2', '3A\\.1\\.2'],
];
const CLASS1_CONTROLS = [
['06', '6'],
['7', '7'],
['08.5', '8\\.5'],
['23.1.2', '23\\.1\\.2'],
];
describe('#4748 — the $((10#$PHASE_INT)) split sites carry a letter suffix into PHASE_N instead of aborting', () => {
const sites = [
// execute-phase.md: plan selection (safe_resume_gate) and the TDD gate a
// few lines below are the same two lines twice; both must be under test.
{ name: 'execute-phase.md', file: EXECUTE_PHASE, anchor: 'PHASE_INT=${PHASE_NUMBER%%', count: 2, input: 'PHASE_NUMBER', output: 'PHASE_N' },
{ name: 'completion-reconciliation.md', file: COMPLETION_RECONCILIATION, anchor: 'SPOT_PHASE_INT=${SPOT_PHASE_NUMBER%%', count: 1, input: 'SPOT_PHASE_NUMBER', output: 'SPOT_PHASE_N' },
{ name: 'tdd.md', file: TDD_REF, anchor: 'PHASE_INT=${PHASE%%', count: 1, input: 'PHASE', output: 'PHASE_N' },
];
for (const site of sites) {
describe(site.name, () => {
const lines = splitLines(fs.readFileSync(site.file, 'utf8'));
const indexes = findAnchoredLineIndexes(lines, site.anchor, site.count);
indexes.forEach((i, n) => {
// The split line and the PHASE_N line directly below it, verbatim.
const splitLine = lines[i].trim();
const nLine = lines[i + 1].trim();
assert.ok(nLine.startsWith(`${site.output}=`), `line after the split must assign ${site.output}: ${nLine}`);
const snippet = ['set -e', splitLine, nLine, `printf '%s' "$${site.output}"`].join('\n');
const label = site.count > 1 ? ` (occurrence ${n + 1})` : '';
for (const [id, expected] of CLASS1_CASES) {
test(`${id} → ${site.output}=${expected} without a shell error${label} (fails before the fix)`, () => {
const r = runBash(snippet, { [site.input]: id });
assert.equal(r.status, 0, `bash exited ${r.status}: ${r.stderr}`);
assert.equal(r.stdout, expected);
});
}
for (const [id, expected] of CLASS1_CONTROLS) {
test(`regression control: ${id} → ${site.output}=${expected}${label}`, () => {
const r = runBash(snippet, { [site.input]: id });
assert.equal(r.status, 0, `bash exited ${r.status}: ${r.stderr}`);
assert.equal(r.stdout, expected);
});
}
test(`the commit-scope ERE built from PHASE_N matches both the padded and the unpadded scope of a letter phase${label}`, () => {
// Each site feeds PHASE_N into `^[a-z]+\((0*${PHASE_N})-(0*${PLAN_N})\):`
// — the #4003 zero-pad-tolerant scope. Prove the value it now yields
// for `03A` matches the two subjects an executor could have written,
// and does NOT match the letter-less phase 3.
const script = [
'set -e',
splitLine,
nLine,
`SCOPE_RE="^[a-z]+\\((0*\${${site.output}})-(0*1)\\):"`,
'for s in "feat(3A-01): x" "feat(03A-1): x"; do printf \'%s\\n\' "$s" | grep -qE "$SCOPE_RE" || { echo "MISS $s"; exit 3; }; done',
'printf \'%s\\n\' "feat(3-01): x" | grep -qE "$SCOPE_RE" && { echo "FALSE-MATCH"; exit 4; }',
'echo OK',
].join('\n');
const r = runBash(script, { [site.input]: '03A' });
assert.equal(r.status, 0, `${r.stdout} ${r.stderr}`);
assert.equal(r.stdout, 'OK');
});
});
});
}
});
describe('#4748 — execute-phase.md resolves the REVIEW.md path from init\'s padded_phase, not a shell re-pad', () => {
const lines = splitLines(fs.readFileSync(EXECUTE_PHASE, 'utf8'));
const [i] = findAnchoredLineIndexes(lines, 'REVIEW_FILE="${PHASE_DIR}/${PADDED}-REVIEW.md"', 1);
const paddedLine = lines[i - 1].trim();
test('the PADDED binding directly above the lookup reads {padded_phase} (fails before the fix)', () => {
// `printf "%02d"` cannot pad `03A` (prints `03`, exits 1) — and cannot
// even re-pad an already-padded `08` (bash reads it as octal, prints
// `00`). `init execute-phase` now emits `padded_phase` through the
// canonical normalizer, so the workflow binds it instead of re-deriving.
assert.ok(paddedLine.startsWith('PADDED='), `line above the lookup must bind PADDED: ${paddedLine}`);
assert.equal(paddedLine, 'PADDED="{padded_phase}"');
});
test('regression control: the lookup line itself is unchanged', () => {
assert.equal(lines[i].trim(), 'REVIEW_FILE="${PHASE_DIR}/${PADDED}-REVIEW.md"');
});
test('composition: the value init emits, substituted into the live lookup lines, resolves the letter phase\'s own REVIEW.md', (t) => {
// The model substitutes `{padded_phase}` from the init JSON, which is
// `normalizePhaseName(phase_number)` (src/init.cts). Do that substitution
// here and run the three live lines against a fixture, so the emitted
// value, the binding, the path construction and the status extraction are
// exercised together — the executable half of a `{template}` site.
const { normalizePhaseName } = require('../gsd-core/bin/lib/phase-id.cjs');
const { createTempDir, cleanup } = require('./helpers.cjs');
const dir = createTempDir();
t.after(() => cleanup(dir));
for (const [id, status] of [['3A', 'clean'], ['8', 'issues'], ['9', 'skipped']]) {
const emitted = normalizePhaseName(id);
fs.writeFileSync(path.join(dir, `${emitted}-REVIEW.md`), `---\nstatus: ${status}\n---\n# review\n`);
const script = [
'set -e',
paddedLine.replace('{padded_phase}', emitted),
lines[i].trim(),
lines[i + 1].trim(),
'printf \'%s %s\' "$PADDED" "$REVIEW_STATUS"',
].join('\n');
assert.ok(lines[i + 1].includes('REVIEW_STATUS='), `line after the lookup must extract REVIEW_STATUS: ${lines[i + 1]}`);
const r = runBash(script, { PHASE_DIR: dir });
assert.equal(r.status, 0, `bash exited ${r.status}: ${r.stderr}`);
assert.equal(r.stdout, `${emitted} ${status}`);
}
});
test('the workflow\'s init parse list names padded_phase, so the binding is not a literal (fails before the fix)', () => {
// A `{field}` token is substituted from the init JSON only for fields the
// workflow tells the model to parse; `phase_number` is on that list and
// `padded_phase` was not (adversarial review, claim 2).
const [p] = findAnchoredLineIndexes(lines, 'Parse JSON for: `executor_model`', 1);
assert.match(lines[p], /`phase_number`, `padded_phase`,/);
});
});
describe('#4748 — autonomous.md --from/--to/--only and plan-review-convergence.md extract the full letter-suffixed phase', () => {
const autonomousText = fs.readFileSync(AUTONOMOUS, 'utf8');
const prcText = fs.readFileSync(PLAN_REVIEW_CONVERGENCE, 'utf8');
const sites = [
{ name: 'autonomous.md --from', pattern: extractGrepPattern(autonomousText, 'FROM_PHASE=$(echo "$ARGUMENTS" | grep -oE'), args: (v) => `--from ${v}`, tail: "| awk '{print $2}'" },
{ name: 'autonomous.md --to', pattern: extractGrepPattern(autonomousText, 'TO_PHASE=$(echo "$ARGUMENTS" | grep -oE'), args: (v) => `--from 1 --to ${v}`, tail: "| awk '{print $2}'" },
{ name: 'autonomous.md --only', pattern: extractGrepPattern(autonomousText, 'ONLY_PHASE=$(echo "$ARGUMENTS" | grep -oE'), args: (v) => `--only ${v} --interactive`, tail: "| awk '{print $2}'" },
{ name: 'plan-review-convergence.md', pattern: extractGrepPattern(prcText, 'PHASE=$(echo "$ARGUMENTS" | grep -oE'), args: (v) => `${v} --codex --max-cycles 3`, tail: '| head -1' },
];
function extract(site, v) {
const script = `echo "$ARGUMENTS" | grep -oE '${site.pattern}' ${site.tail}`;
return runBash(script, { ARGUMENTS: site.args(v) }).stdout;
}
for (const site of sites) {
describe(site.name, () => {
// Before the fix `[0-9]+\.?[0-9]*` stops at the letter: `12A` → `12`,
// silently targeting a different phase. `23.1.2` → `23.1` is the same
// truncation one axis over (#4568's class in a spelling neither lint saw).
for (const v of ['12A', '3A', '23A.1.2', '23.1.2']) {
test(`${v} extracts ${v}, not a truncated prefix (fails before the fix)`, () => {
const got = extract(site, v);
assert.equal(got, v);
assert.equal(CANONICAL_ANCHORED.test(got), true);
});
}
for (const v of ['6', '36.14']) {
test(`regression control: ${v} extracts ${v}`, () => {
assert.equal(extract(site, v), v);
});
}
});
}
test('autonomous.md: the three flags extract independently from one argument string', () => {
const script = [
`FROM_PHASE=$(echo "$ARGUMENTS" | grep -oE '${sites[0].pattern}' | awk '{print $2}')`,
`TO_PHASE=$(echo "$ARGUMENTS" | grep -oE '${sites[1].pattern}' | awk '{print $2}')`,
'printf \'%s %s\' "$FROM_PHASE" "$TO_PHASE"',
].join('\n');
assert.equal(runBash(script, { ARGUMENTS: '--from 3A --to 5B --max-cycles 2' }).stdout, '3A 5B');
});
});
describe('#4748 — phase-argument-parsing.md\'s legacy normalizer pads a letter-suffixed id instead of leaving it alone', () => {
const lines = splitLines(fs.readFileSync(PHASE_ARGUMENT_PARSING, 'utf8'));
const [start] = findAnchoredLineIndexes(lines, '# Normalize phase number', 1);
let end = start;
while (end < lines.length && lines[end].trim() !== 'fi') end++;
assert.ok(end < lines.length, 'normalizer block must close with `fi`');
const block = lines.slice(start, end + 1).join('\n');
function normalize(v) {
return runBash(`set -e\n${block}\nprintf '%s' "$PHASE"`, { PHASE: v });
}
// Before the fix neither branch matches a letter id, so `12A` passes through
// unpadded and `3A` is never zero-padded to the `03A` a directory carries.
for (const [input, expected] of [['3A', '03A'], ['12A', '12A'], ['3A.1', '03A.1'], ['23A.1.2', '23A.1.2']]) {
test(`${input} → ${expected} (fails before the fix)`, () => {
const r = normalize(input);
assert.equal(r.status, 0, `bash exited ${r.status}: ${r.stderr}`);
assert.equal(r.stdout, expected);
assert.equal(CANONICAL_ANCHORED.test(r.stdout), true);
});
}
// `08` is the octal trap: `printf "%02d" 08` is an invalid octal number in
// bash (exit 1, prints `00`), so the old integer branch mangled any
// already-padded id it was handed. `23.1.2` matched neither old branch and
// passed through unchanged — the N-segment axis was silently unpadded.
for (const [input, expected] of [['08', '08'], ['23.1.2', '23.1.2']]) {
test(`${input} → ${expected} without a shell error (fails before the fix)`, () => {
const r = normalize(input);
assert.equal(r.status, 0, `bash exited ${r.status}: ${r.stderr}`);
assert.equal(r.stdout, expected);
});
}
for (const [input, expected] of [['8', '08'], ['2.1', '02.1'], ['36.14', '36.14']]) {
test(`regression control: ${input} → ${expected}`, () => {
const r = normalize(input);
assert.equal(r.status, 0, `bash exited ${r.status}: ${r.stderr}`);
assert.equal(r.stdout, expected);
});
}
test('a non-canonical value passes through untouched (the normalizer is not a validator)', () => {
const r = normalize('AUTH-101');
assert.equal(r.status, 0);
assert.equal(r.stdout, 'AUTH-101');
});
});